Part A: Guided lab#
Objective#
By the end of the lab you will have written the core of a RAG retriever and measured it:
chunk_text: split documents into overlapping word windows.cosine_similarity: the similarity measure behind vector search.search: find the top-k chunks for a question with one matrix-vector product.build_prompt: turn retrieved chunks into a grounded prompt for an LLM.evaluate_retrieval: compute hit rate@k and recall@k against hand-labelled questions.
You will also compare an embedding model with TF-IDF keyword vectors and explain the difference using your own results.
Requirements#
| Item | Detail |
|---|---|
| Time | 60–90 minutes (timings per step below) |
| Python | 3.10 or newer (tested on 3.13.5) |
| Hardware | Any laptop; no GPU needed |
| Network | Needed once to install packages and download the all-MiniLM-L6-v2 model. If the download is blocked, the code falls back to TF-IDF automatically and the lab still works. |
| Accounts / API keys | None. Nothing in this lab costs money. |
| Knowledge | Python functions, lists, dicts, f-strings. NumPy basics help but every line is explained. |
Package versions used to produce the expected outputs:
| Package | Version |
|---|---|
| Python | 3.13.5 |
| numpy | 2.5.3 |
| scikit-learn | 1.9.1 |
| sentence-transformers | 6.1.0 |
| torch | 2.14.0+cpu |
| transformers | 5.17.0 |
The scenario#
Kestrelwood Systems, a fictional company, wants an assistant that answers staff questions about HR and IT policies. You have ten short policy documents (annual leave, sick leave, hybrid work, expenses, travel, passwords, laptops, learning budget, parental leave, data and AI tool usage) and ten questions an HR colleague labelled with the correct document. Every name, number and address in the data is made up.
Step 0: Set up (10 minutes)#
mkdir rag-lab && cd rag-lab
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# Optional on Linux: CPU-only PyTorch avoids a large CUDA download
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install numpy scikit-learn sentence-transformers
Create two files in the rag-lab folder:
kestrelwood_data.py: the corpus and labelled questions (copy from the block below; do not edit).rag_lab_starter.py: the starter code from the next section.
kestrelwood_data.py (click to expand and copy)
"""Kestrelwood Systems sample corpus (FICTIONAL company, made-up policies) and a hand-labelled
evaluation set. Identical to the data in rag_from_scratch.py."""
CORPUS = [
{
"doc_id": "HR-01",
"title": "Annual leave policy",
"text": (
"Every full-time employee at Kestrelwood Systems receives 24 days of paid annual "
"leave per calendar year. Leave accrues at 2 days per month. You can carry forward "
"a maximum of 10 unused days into the next calendar year; any balance above 10 days "
"lapses on 31 December. Apply for leave in the PeoplePortal at least 5 working days "
"in advance. Your reporting manager approves or rejects the request within 2 working days."
),
},
{
"doc_id": "HR-02",
"title": "Sick leave policy",
"text": (
"Employees receive 12 days of paid sick leave per calendar year. Sick leave cannot be "
"carried forward. Inform your manager before 10:00 AM on the day you are unwell. "
"If you are absent for more than 2 consecutive days, upload a medical certificate "
"from a registered doctor to the PeoplePortal within 3 days of returning to work."
),
},
{
"doc_id": "HR-03",
"title": "Hybrid and remote work",
"text": (
"Kestrelwood follows a hybrid model. You may work remotely up to 3 days per week. "
"Wednesday is the team anchor day and everyone works from the office. Core "
"collaboration hours are 11:00 to 16:00 IST, when you must be reachable on chat. "
"When working outside the office, always connect through the company VPN."
),
},
{
"doc_id": "FIN-01",
"title": "Expense reimbursement",
"text": (
"Submit expense claims in ExpenseDesk within 30 days of spending the money. Attach "
"an itemised receipt for every line item; card statements are not accepted as receipts. "
"Approved claims are paid with the next monthly payroll. Any single claim above "
"INR 25,000 needs approval from your department head in addition to your manager."
),
},
{
"doc_id": "FIN-02",
"title": "Business travel policy",
"text": (
"Book all business travel through the internal travel desk at least 7 days before the "
"trip. Flights shorter than 6 hours must be booked in economy class. The hotel limit "
"is INR 6,000 per night in metro cities and INR 4,000 per night elsewhere. A daily "
"allowance of INR 1,500 covers meals and local transport."
),
},
{
"doc_id": "IT-01",
"title": "Passwords and account security",
"text": (
"Passwords must be at least 14 characters long. Multi-factor authentication is "
"mandatory for every company account. If you forget your password, reset it yourself "
"at the self-service portal id.kestrelwood.example using your registered phone. "
"Never share one-time passcodes with anyone. The IT team will never ask for your password."
),
},
{
"doc_id": "IT-02",
"title": "Laptops and devices",
"text": (
"Every new joiner receives a company laptop on day one. All laptops use full-disk "
"encryption and are managed centrally. If your laptop or phone is lost or stolen, "
"report it to the IT helpdesk on extension 4040 within 2 hours so the device can be "
"locked and wiped remotely. Personal devices may access email only through the managed browser."
),
},
{
"doc_id": "LND-01",
"title": "Learning and certification budget",
"text": (
"Each employee has a learning budget of INR 50,000 per financial year for courses, "
"books and professional certification exams. Get written pre-approval from your manager "
"before you pay. Claim the cost through ExpenseDesk with the invoice and proof of completion. "
"If you leave the company within 12 months of a reimbursement, you repay 50 percent of it."
),
},
{
"doc_id": "HR-04",
"title": "Parental leave",
"text": (
"Birth mothers receive 26 weeks of paid maternity leave. Fathers, partners and "
"adoptive parents receive 4 weeks of paid parental leave, which must be taken within "
"6 months of the birth or adoption. After parental leave, employees can request a "
"phased return with reduced hours for up to 8 weeks."
),
},
{
"doc_id": "SEC-01",
"title": "Data classification and AI tool usage",
"text": (
"Company information is classified as Public, Internal, Confidential or Restricted. "
"Customer personal data is always Restricted. Never paste Confidential or Restricted "
"data into external AI chatbots or public websites. Use only the approved internal "
"assistant for work involving customer data, and report any accidental disclosure to "
"security@kestrelwood.example immediately."
),
},
]
# Hand-labelled evaluation set: question -> set of doc_ids that contain the answer.
EVAL_SET = [
{"question": "How many vacation days do I get each year?", "relevant": {"HR-01"}},
{"question": "Can I carry unused leave into next year?", "relevant": {"HR-01"}},
{"question": "Do I need a doctor's note if I am ill for three days?", "relevant": {"HR-02"}},
{"question": "I forgot my password. How do I reset it?", "relevant": {"IT-01"}},
{"question": "My laptop was stolen at the airport. What should I do?", "relevant": {"IT-02"}},
{"question": "What is the hotel limit when I travel to Mumbai?", "relevant": {"FIN-02"}},
{"question": "Can I use ChatGPT to summarise a customer's complaint?", "relevant": {"SEC-01"}},
{"question": "How much time off do new fathers get?", "relevant": {"HR-04"}},
{"question": "Will the company pay for my AWS certification exam?", "relevant": {"LND-01"}},
{
"question": "How do I book flights for a client visit and claim the costs afterwards?",
"relevant": {"FIN-02", "FIN-01"},
},
]
Starter code: rag_lab_starter.py#
Five functions contain a # TODO and raise NotImplementedError. Everything else (build_chunks, the Embedder class, build_index) is complete; read it, but you don't need to change it.
Download: full code bundle (rag-tutorial-code.zip), which includes the lab starter and solution.
"""RAG lab - starter code. Fill in the five TODOs, then run: python rag_lab_starter.py"""
import numpy as np
from kestrelwood_data import CORPUS, EVAL_SET
def chunk_text(text: str, chunk_size: int = 50, overlap: int = 10) -> list[str]:
"""Split text into chunks of `chunk_size` words; consecutive chunks share `overlap` words."""
if chunk_size <= 0:
raise ValueError("chunk_size must be positive")
if not 0 <= overlap < chunk_size:
raise ValueError("overlap must be >= 0 and smaller than chunk_size")
words = text.split()
step = chunk_size - overlap
chunks = []
# TODO 1: loop over start positions 0, step, 2*step, ...
# append " ".join(words[start:start + chunk_size]) to chunks
# and stop once a chunk reaches the end of the text.
raise NotImplementedError("TODO 1: chunk_text")
return chunks
def build_chunks(corpus: list[dict], chunk_size: int = 50, overlap: int = 10) -> list[dict]:
"""Chunk every document and keep metadata (doc_id, title) with each chunk."""
chunks = []
for doc in corpus:
for i, piece in enumerate(chunk_text(doc["text"], chunk_size, overlap)):
chunks.append({
"chunk_id": f"{doc['doc_id']}#{i}",
"doc_id": doc["doc_id"],
"title": doc["title"],
"text": piece,
})
return chunks
class Embedder:
"""Turns text into L2-normalised vectors.
backend="auto" -> try sentence-transformers all-MiniLM-L6-v2, fall back to TF-IDF
backend="minilm" -> sentence-transformers only
backend="tfidf" -> scikit-learn TF-IDF only (no download needed)
"""
MODEL_NAME = "sentence-transformers/all-MiniLM-L6-v2"
def __init__(self, backend: str = "auto"):
self.backend = None
self.model = None
if backend in ("auto", "minilm"):
try:
from sentence_transformers import SentenceTransformer
self.model = SentenceTransformer(self.MODEL_NAME)
self.backend = "minilm"
except Exception as exc: # no package, no network, blocked download...
if backend == "minilm":
raise
print(f"[info] sentence-transformers unavailable ({exc!r}); using TF-IDF.")
if self.backend is None:
from sklearn.feature_extraction.text import TfidfVectorizer
self.model = TfidfVectorizer(stop_words="english", ngram_range=(1, 2))
self.backend = "tfidf"
self._fitted = self.backend == "minilm"
def fit(self, texts: list[str]) -> "Embedder":
"""TF-IDF must learn its vocabulary from the corpus; MiniLM is already trained."""
if self.backend == "tfidf":
self.model.fit(texts)
self._fitted = True
return self
def encode(self, texts: list[str]) -> np.ndarray:
if not self._fitted:
raise RuntimeError("Call fit() on the corpus before encode() when using TF-IDF.")
if self.backend == "minilm":
vectors = self.model.encode(texts, convert_to_numpy=True)
else:
vectors = self.model.transform(texts).toarray()
vectors = vectors.astype(np.float32)
norms = np.linalg.norm(vectors, axis=1, keepdims=True)
norms[norms == 0] = 1.0 # avoid division by zero for empty vectors
return vectors / norms # unit length -> dot product == cosine similarity
def cosine_similarity(a: np.ndarray, b: np.ndarray) -> float:
"""cos(theta) = (a . b) / (||a|| * ||b||)"""
# TODO 2: use np.dot and np.linalg.norm. Return a Python float.
raise NotImplementedError("TODO 2: cosine_similarity")
def build_index(chunks: list[dict], embedder: Embedder) -> np.ndarray:
"""Embed every chunk. Row i of the returned matrix is the vector for chunks[i]."""
texts = [c["text"] for c in chunks]
embedder.fit(texts)
return embedder.encode(texts)
def search(query: str, index: np.ndarray, chunks: list[dict], embedder: Embedder, k: int = 3) -> list[dict]:
"""Return the k chunks most similar to the query, best first."""
q = embedder.encode([query])[0]
# TODO 3: compute one score per chunk with a single matrix-vector product (index @ q),
# take the indices of the k highest scores (np.argsort on -scores),
# and return [{**chunks[i], "score": float(scores[i])} for i in top].
raise NotImplementedError("TODO 3: search")
SYSTEM_PROMPT = (
"You are the HR and IT help assistant for Kestrelwood Systems. "
"Answer ONLY from the context provided. Cite the source id in square brackets, e.g. [HR-01]. "
"If the context does not contain the answer, reply exactly: "
"\"I don't know based on the company documents.\""
)
def build_prompt(question: str, retrieved: list[dict]) -> str:
"""Put the retrieved chunks and the question into one user message."""
# TODO 4: build a context string where each chunk looks like
# "[DOC-ID] Title\n<chunk text>", separated by a blank line, then return
# f"Context:\n{context}\n\nQuestion: {question}\nAnswer:"
raise NotImplementedError("TODO 4: build_prompt")
def evaluate_retrieval(eval_set: list[dict], index: np.ndarray, chunks: list[dict],
embedder: Embedder, k: int = 3, verbose: bool = False) -> dict:
"""Hit rate@k: share of questions with at least one relevant doc in the top k.
Recall@k: average share of each question's relevant docs found in the top k."""
hits, recalls = [], []
for item in eval_set:
results = search(item["question"], index, chunks, embedder, k=k)
# TODO 5: found = the set of doc_ids in results that are also in item["relevant"]
# append 1.0 or 0.0 to hits, and len(found) / len(item["relevant"]) to recalls
raise NotImplementedError("TODO 5: evaluate_retrieval")
return {"k": k, "hit_rate": float(np.mean(hits)), "recall": float(np.mean(recalls))}
if __name__ == "__main__":
print("Checkpoint 1:", chunk_text("a b c d e f g h i j", chunk_size=4, overlap=1))
chunks = build_chunks(CORPUS, chunk_size=50, overlap=10)
print("Checkpoint 2:", len(chunks), "chunks; first id =", chunks[0]["chunk_id"])
print("Checkpoint 3:", round(cosine_similarity(np.array([1, 2, 2]), np.array([2, 1, 2])), 3),
round(cosine_similarity(np.array([1, 2, 2]), np.array([0, 6, 0])), 3))
embedder = Embedder("auto")
index = build_index(chunks, embedder)
print("Checkpoint 4:", embedder.backend, index.shape)
print("Checkpoint 5:")
for r in search("Can I carry unused leave into next year?", index, chunks, embedder, k=3):
print(f" {r['score']:.3f} {r['chunk_id']:<9} {r['title']}")
print("Checkpoint 6:")
top2 = search("Do I need a doctor's note if I am ill for three days?", index, chunks, embedder, k=2)
print(build_prompt("Do I need a doctor's note if I am ill for three days?", top2))
print("Checkpoint 7:")
for k in (1, 3):
m = evaluate_retrieval(EVAL_SET, index, chunks, embedder, k=k)
print(f" k={k}: hit_rate={m['hit_rate']:.2f} recall={m['recall']:.2f}")
Run it now to confirm your setup:
python rag_lab_starter.py
Expected output (the traceback ends with these lines; your folder path will differ):
File ".../rag-lab/rag_lab_starter.py", line 19, in chunk_text
raise NotImplementedError("TODO 1: chunk_text")
NotImplementedError: TODO 1: chunk_text
That error is correct: it means Python, the data file and the imports all work. Now fill in the TODOs in order. After each one, run python rag_lab_starter.py again. Each checkpoint prints once its functions work, and the run stops at the next unfinished TODO.
Step 1: Chunking with overlap, TODO 1 (15 minutes)#
Why: an embedding squeezes a whole text into one vector. Long texts produce vague vectors, and embedding models truncate long input (all-MiniLM-L6-v2 truncates beyond 256 word pieces). Overlap repeats a few words at each boundary so a fact cut in half still appears whole in one chunk.
Your task: loop over start positions 0, step, 2*step, ... where step = chunk_size - overlap. Take words[start:start + chunk_size], join with spaces, append. Stop as soon as a chunk reaches the end of the text, otherwise you'll create a last chunk that's entirely overlap.
Hint: for start in range(0, len(words), step): then if start + chunk_size >= len(words): break after appending.
Checkpoint 1 and 2, expected output:
Checkpoint 1: ['a b c d', 'd e f g', 'g h i j']
Checkpoint 2: 17 chunks; first id = HR-01#0
Check the first line by hand: 10 letters, window 4, overlap 1, so the window moves by 3 and each chunk starts with the last letter of the previous one.
Step 2: Cosine similarity, TODO 2 (5 minutes)#
Why: search ranks chunks by how closely their vectors point in the same direction as the question vector. Cosine similarity measures that direction and ignores length:
where \(\mathbf{a}\cdot\mathbf{b}\) is the dot product (multiply matching positions and add), and \(\lVert\mathbf{a}\rVert\) is the vector length (square root of the sum of squares).
Your task: one line with np.dot and np.linalg.norm, wrapped in float(...).
Checkpoint 3, expected output:
Checkpoint 3: 0.889 0.667
Verify by hand: for \((1,2,2)\) and \((2,1,2)\), the dot product is \(8\), both lengths are \(3\), so \(8/9 = 0.889\). For \((1,2,2)\) and \((0,6,0)\), the dot product is \(12\) (bigger!) but the lengths are \(3\) and \(6\), so \(12/18 = 0.667\). The longer vector wins on dot product and loses on cosine.
Step 3: Embeddings (5 minutes, no TODO)#
Read the Embedder class. Note three things:
- With
backend="auto"it triesall-MiniLM-L6-v2(384 numbers per text) and falls back to TF-IDF if that fails. - TF-IDF must
fit()on your chunks to learn a vocabulary; MiniLM is already trained. encode()divides every vector by its length, so each vector has length 1 and a plain dot product equals cosine similarity.
Checkpoint 4, expected output:
Checkpoint 4: minilm (17, 384)
17 chunks, 384 dimensions each. If you see tfidf (17, 525) instead, the model download failed and the fallback kicked in (see Troubleshooting). You can continue; your later numbers will match the TF-IDF column in Step 7.
Step 4: Search, TODO 3 (10 minutes)#
Why: with unit-length vectors, index @ q computes the cosine similarity between the question and every chunk in one operation: row \(j\) of the index times \(\mathbf{q}\) is the score of chunk \(j\).
Your task: compute scores = index @ q, get top = np.argsort(-scores)[:k] (the minus sign because argsort sorts ascending), and return a list of chunk dicts with a "score" key added.
Checkpoint 5, expected output:
Checkpoint 5:
0.505 HR-01#0 Annual leave policy
0.426 HR-01#1 Annual leave policy
0.426 HR-02#0 Sick leave policy
The best chunk is the first chunk of the annual leave policy, which contains the carry-forward rule. Your third decimal place may differ slightly on other hardware or library versions; the order should not.
Step 5: Build the prompt, TODO 4 (10 minutes)#
Why: this string is exactly what the LLM will read. Labelling each chunk with its document id lets the model cite sources, and the SYSTEM_PROMPT (already in the starter) tells it to answer only from this context and to reply with a fixed sentence when the answer isn't there.
Your task: for each retrieved chunk, format "[DOC-ID] Title\n<chunk text>", join them with a blank line ("\n\n"), and return f"Context:\n{context}\n\nQuestion: {question}\nAnswer:".
Checkpoint 6, expected output:
Checkpoint 6:
Context:
[HR-02] Sick leave policy
Employees receive 12 days of paid sick leave per calendar year. Sick leave cannot be carried forward. Inform your manager before 10:00 AM on the day you are unwell. If you are absent for more than 2 consecutive days, upload a medical certificate from a registered doctor to the PeoplePortal
[HR-02] Sick leave policy
a medical certificate from a registered doctor to the PeoplePortal within 3 days of returning to work.
Question: Do I need a doctor's note if I am ill for three days?
Answer:
Look closely: both chunks are from HR-02, and the phrase "a medical certificate from a registered doctor to the PeoplePortal" appears twice. That's the overlap from Step 1. It helped retrieval (the second chunk holds the "within 3 days" detail), but it also wastes prompt space. Exercise 2 at the end asks you to fix this.
Step 6: Evaluate retrieval, TODO 5 (10 minutes)#
Why: trying three questions by hand and saying "looks good" is not a measurement. With labelled questions you get numbers you can compare after every change.
- Hit rate@k: fraction of questions where at least one relevant document appears in the top k.
- Recall@k: for each question, the fraction of its relevant documents found in the top k, then averaged.
Your task: found = {r["doc_id"] for r in results} & item["relevant"], then append 1.0 if found else 0.0 to hits and len(found) / len(item["relevant"]) to recalls.
Checkpoint 7, expected output (your full run should now finish):
Checkpoint 7:
k=1: hit_rate=1.00 recall=0.95
k=3: hit_rate=1.00 recall=1.00
Why is recall@1 below hit rate@1? One question ("How do I book flights for a client visit and claim the costs afterwards?") needs two documents, FIN-02 and FIN-01. One chunk can only come from one document, so at k=1 that question scores recall 0.5, and the average becomes (9 × 1 + 0.5) / 10 = 0.95.
Step 7: Experiment and explain (15–20 minutes)#
7a. Keywords versus meaning. In rag_lab_starter.py, change Embedder("auto") to Embedder("tfidf") and run again. Expected output:
Checkpoint 1: ['a b c d', 'd e f g', 'g h i j']
Checkpoint 2: 17 chunks; first id = HR-01#0
Checkpoint 3: 0.889 0.667
Checkpoint 4: tfidf (17, 525)
Checkpoint 5:
0.317 HR-01#0 Annual leave policy
0.110 HR-02#0 Sick leave policy
0.081 LND-01#0 Learning and certification budget
Checkpoint 6:
Context:
[HR-02] Sick leave policy
a medical certificate from a registered doctor to the PeoplePortal within 3 days of returning to work.
[HR-02] Sick leave policy
Employees receive 12 days of paid sick leave per calendar year. Sick leave cannot be carried forward. Inform your manager before 10:00 AM on the day you are unwell. If you are absent for more than 2 consecutive days, upload a medical certificate from a registered doctor to the PeoplePortal
Question: Do I need a doctor's note if I am ill for three days?
Answer:
Checkpoint 7:
k=1: hit_rate=0.90 recall=0.85
k=3: hit_rate=1.00 recall=1.00
Compare with your MiniLM run and write down answers to these questions:
- Hit rate@1 fell from 1.00 to 0.90. Temporarily call
evaluate_retrieval(..., k=1, verbose=True)to find the question that missed. (You should find "How much time off do new fathers get?", which TF-IDF sends to the laptop policy. Why might the word "new" do that? Look at the first sentence of IT-02.) - In Checkpoint 6 the order of the two HR-02 chunks flipped. Which words does the short chunk share with the question, and why might a shorter chunk score higher than a longer one with the same matching words? (Hint: every vector is normalised to length 1.)
- Scores are much lower with TF-IDF (0.317 against 0.505 for the top result). Does that mean TF-IDF is "less confident"? (Hint: can you compare scores from two different vector spaces?)
7b. Chunk size. Switch back to Embedder("auto"). In the __main__ block, change build_chunks(CORPUS, chunk_size=50, overlap=10): try chunk_size=12, overlap=0, then chunk_size=12, overlap=4. Record hit rate@1 each time. For reference, the tutorial's runs gave hit@1 = 0.90 for both 12-word settings, against 1.00 at 50/10. Then print the top chunk for "How much time off do new fathers get?" in each setting and check whether it actually contains the answer ("4 weeks").
Stretch goals (optional)#
- Generation. The tutorial shows
generate_with_bedrock(), which sends yourbuild_promptoutput to an Amazon Bedrock model through the Converse API. Bedrock is billed per input and output token. Only try this with your instructor's approval and an AWS account you are allowed to use. Any chat LLM can take its place: the prompt is just a string. - Vector database. Load your chunks and vectors into FAISS (
IndexFlatIP) or Chroma (withconfiguration={"hnsw": {"space": "cosine"}}) and confirm you get the same top-3 as yoursearch(). The tutorial has working code.
Full solution: rag_lab_solution.py#
Try the TODOs yourself before reading this. The solution is identical to the functions in the tutorial's rag_from_scratch.py.
Download: full code bundle (rag-tutorial-code.zip) · main script only (rag_from_scratch.py)
Show the full solution
"""RAG lab - full solution. Same function names as rag_from_scratch.py."""
import numpy as np
from kestrelwood_data import CORPUS, EVAL_SET
def chunk_text(text: str, chunk_size: int = 50, overlap: int = 10) -> list[str]:
"""Split text into chunks of `chunk_size` words; consecutive chunks share `overlap` words."""
if chunk_size <= 0:
raise ValueError("chunk_size must be positive")
if not 0 <= overlap < chunk_size:
raise ValueError("overlap must be >= 0 and smaller than chunk_size")
words = text.split()
step = chunk_size - overlap
chunks = []
for start in range(0, len(words), step):
chunks.append(" ".join(words[start:start + chunk_size]))
if start + chunk_size >= len(words):
break
return chunks
def build_chunks(corpus: list[dict], chunk_size: int = 50, overlap: int = 10) -> list[dict]:
"""Chunk every document and keep metadata (doc_id, title) with each chunk."""
chunks = []
for doc in corpus:
for i, piece in enumerate(chunk_text(doc["text"], chunk_size, overlap)):
chunks.append({
"chunk_id": f"{doc['doc_id']}#{i}",
"doc_id": doc["doc_id"],
"title": doc["title"],
"text": piece,
})
return chunks
class Embedder:
"""Turns text into L2-normalised vectors.
backend="auto" -> try sentence-transformers all-MiniLM-L6-v2, fall back to TF-IDF
backend="minilm" -> sentence-transformers only
backend="tfidf" -> scikit-learn TF-IDF only (no download needed)
"""
MODEL_NAME = "sentence-transformers/all-MiniLM-L6-v2"
def __init__(self, backend: str = "auto"):
self.backend = None
self.model = None
if backend in ("auto", "minilm"):
try:
from sentence_transformers import SentenceTransformer
self.model = SentenceTransformer(self.MODEL_NAME)
self.backend = "minilm"
except Exception as exc: # no package, no network, blocked download...
if backend == "minilm":
raise
print(f"[info] sentence-transformers unavailable ({exc!r}); using TF-IDF.")
if self.backend is None:
from sklearn.feature_extraction.text import TfidfVectorizer
self.model = TfidfVectorizer(stop_words="english", ngram_range=(1, 2))
self.backend = "tfidf"
self._fitted = self.backend == "minilm"
def fit(self, texts: list[str]) -> "Embedder":
"""TF-IDF must learn its vocabulary from the corpus; MiniLM is already trained."""
if self.backend == "tfidf":
self.model.fit(texts)
self._fitted = True
return self
def encode(self, texts: list[str]) -> np.ndarray:
if not self._fitted:
raise RuntimeError("Call fit() on the corpus before encode() when using TF-IDF.")
if self.backend == "minilm":
vectors = self.model.encode(texts, convert_to_numpy=True)
else:
vectors = self.model.transform(texts).toarray()
vectors = vectors.astype(np.float32)
norms = np.linalg.norm(vectors, axis=1, keepdims=True)
norms[norms == 0] = 1.0 # avoid division by zero for empty vectors
return vectors / norms # unit length -> dot product == cosine similarity
def cosine_similarity(a: np.ndarray, b: np.ndarray) -> float:
"""cos(theta) = (a . b) / (||a|| * ||b||)"""
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
def build_index(chunks: list[dict], embedder: Embedder) -> np.ndarray:
"""Embed every chunk. Row i of the returned matrix is the vector for chunks[i]."""
texts = [c["text"] for c in chunks]
embedder.fit(texts)
return embedder.encode(texts)
def search(query: str, index: np.ndarray, chunks: list[dict], embedder: Embedder, k: int = 3) -> list[dict]:
"""Return the k chunks most similar to the query, best first."""
q = embedder.encode([query])[0]
scores = index @ q # one dot product per chunk (all vectors are unit length)
top = np.argsort(-scores)[:k] # indices of the k highest scores
return [{**chunks[i], "score": float(scores[i])} for i in top]
SYSTEM_PROMPT = (
"You are the HR and IT help assistant for Kestrelwood Systems. "
"Answer ONLY from the context provided. Cite the source id in square brackets, e.g. [HR-01]. "
"If the context does not contain the answer, reply exactly: "
"\"I don't know based on the company documents.\""
)
def build_prompt(question: str, retrieved: list[dict]) -> str:
"""Put the retrieved chunks and the question into one user message."""
context = "\n\n".join(f"[{r['doc_id']}] {r['title']}\n{r['text']}" for r in retrieved)
return f"Context:\n{context}\n\nQuestion: {question}\nAnswer:"
def evaluate_retrieval(eval_set: list[dict], index: np.ndarray, chunks: list[dict],
embedder: Embedder, k: int = 3, verbose: bool = False) -> dict:
"""Hit rate@k: share of questions with at least one relevant doc in the top k.
Recall@k: average share of each question's relevant docs found in the top k."""
hits, recalls = [], []
for item in eval_set:
results = search(item["question"], index, chunks, embedder, k=k)
found = {r["doc_id"] for r in results} & item["relevant"]
hits.append(1.0 if found else 0.0)
recalls.append(len(found) / len(item["relevant"]))
if verbose:
got = ", ".join(r["doc_id"] for r in results)
mark = "HIT " if found else "MISS"
print(f" {mark} want={sorted(item['relevant'])} got=[{got}] {item['question']}")
return {"k": k, "hit_rate": float(np.mean(hits)), "recall": float(np.mean(recalls))}
if __name__ == "__main__":
print("Checkpoint 1:", chunk_text("a b c d e f g h i j", chunk_size=4, overlap=1))
chunks = build_chunks(CORPUS, chunk_size=50, overlap=10)
print("Checkpoint 2:", len(chunks), "chunks; first id =", chunks[0]["chunk_id"])
print("Checkpoint 3:", round(cosine_similarity(np.array([1, 2, 2]), np.array([2, 1, 2])), 3),
round(cosine_similarity(np.array([1, 2, 2]), np.array([0, 6, 0])), 3))
embedder = Embedder("auto")
index = build_index(chunks, embedder)
print("Checkpoint 4:", embedder.backend, index.shape)
print("Checkpoint 5:")
for r in search("Can I carry unused leave into next year?", index, chunks, embedder, k=3):
print(f" {r['score']:.3f} {r['chunk_id']:<9} {r['title']}")
print("Checkpoint 6:")
top2 = search("Do I need a doctor's note if I am ill for three days?", index, chunks, embedder, k=2)
print(build_prompt("Do I need a doctor's note if I am ill for three days?", top2))
print("Checkpoint 7:")
for k in (1, 3):
m = evaluate_retrieval(EVAL_SET, index, chunks, embedder, k=k)
print(f" k={k}: hit_rate={m['hit_rate']:.2f} recall={m['recall']:.2f}")
Expected output of the full solution (python rag_lab_solution.py):
Checkpoint 1: ['a b c d', 'd e f g', 'g h i j']
Checkpoint 2: 17 chunks; first id = HR-01#0
Checkpoint 3: 0.889 0.667
Checkpoint 4: minilm (17, 384)
Checkpoint 5:
0.505 HR-01#0 Annual leave policy
0.426 HR-01#1 Annual leave policy
0.426 HR-02#0 Sick leave policy
Checkpoint 6:
Context:
[HR-02] Sick leave policy
Employees receive 12 days of paid sick leave per calendar year. Sick leave cannot be carried forward. Inform your manager before 10:00 AM on the day you are unwell. If you are absent for more than 2 consecutive days, upload a medical certificate from a registered doctor to the PeoplePortal
[HR-02] Sick leave policy
a medical certificate from a registered doctor to the PeoplePortal within 3 days of returning to work.
Question: Do I need a doctor's note if I am ill for three days?
Answer:
Checkpoint 7:
k=1: hit_rate=1.00 recall=0.95
k=3: hit_rate=1.00 recall=1.00
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
NotImplementedError: TODO n: ... |
You haven't finished TODO n yet. | Expected; complete that function. |
ModuleNotFoundError: No module named 'kestrelwood_data' |
kestrelwood_data.py is missing or not in the folder you run from. |
Put both files in the same folder and run python rag_lab_starter.py from that folder. |
ModuleNotFoundError: No module named 'sentence_transformers' (or numpy, sklearn) |
Packages installed in a different environment. | Activate the venv (source .venv/bin/activate) and re-run the pip install line. |
[info] sentence-transformers unavailable (OSError("We couldn't connect to 'https://huggingface.co' to load the files, and couldn't find them in the cached files. ...")); using TF-IDF. |
The model download is blocked (offline, proxy or firewall). | Nothing breaks: the lab continues with TF-IDF and your numbers should match Step 7a. To use MiniLM, run once with internet access so the model is cached. |
ValueError: matmul: Input operand 1 has a mismatch in its core dimension 0, ... (size 1 is different from 384) |
In search, you used embedder.encode([query]) without [0], so q has shape (1, 384) instead of (384,). |
Use q = embedder.encode([query])[0]. |
ValueError: matmul: ... (size 525 is different from 384) |
The index and the query were encoded by different embedders (here MiniLM for the index, TF-IDF for the query). | Build the index and encode queries with the same Embedder object. Vectors from different models are not comparable. |
RuntimeError: Call fit() on the corpus before encode() when using TF-IDF. |
You called encode() on a new TF-IDF embedder without fitting it. |
Use build_index(), which calls fit() first. |
ValueError: overlap must be >= 0 and smaller than chunk_size |
For example chunk_size=3, overlap=3: the window would never move. |
Keep overlap smaller than chunk_size. |
| Checkpoint 5 lists the least relevant chunks, with scores increasing down the list. | np.argsort(scores) without the minus sign sorts ascending. |
Use np.argsort(-scores)[:k]. |
Checkpoint 1 prints an extra short chunk at the end, such as 'j'. |
The loop didn't stop when a chunk reached the end of the text. | Add if start + chunk_size >= len(words): break after appending. |
| Your scores differ from the expected output in the third decimal place. | Different hardware or library versions. | Fine, as long as the ranking and the metrics match. |
pip install sentence-transformers downloads a very large PyTorch build. |
pip picked a GPU (CUDA) build of PyTorch. | Install the CPU wheel first: pip install torch --index-url https://download.pytorch.org/whl/cpu. |
Exit ticket (5 minutes)#
Write one sentence for each:
- Why did the embedding model find the ChatGPT question's answer when the word "ChatGPT" is not in any document?
- Why is hit rate@1 alone not enough to prove your assistant would answer correctly? (Think about the 12-word, no-overlap chunk for the "new fathers" question.)
- What is the first thing you would print when a RAG answer is wrong?
Exercises after class#
- Add mean reciprocal rank (MRR) to
evaluate_retrieval: for each question, 1 / rank of the first relevant document, or 0 if none, averaged over questions. - Remove the duplicated overlap text from the prompt: when two retrieved chunks come from the same document, merge them into one context block. Compare prompt length before and after on Checkpoint 6.
- Add five questions the documents can't answer (for example "What is Kestrelwood's share price?"). Print the top score for each. Can you find one similarity threshold that separates them from the ten answerable questions? What does that tell you about relying on thresholds?
Part B: RAG cheat sheet (one page)#
Pipeline. Indexing, when documents change: ingest → chunk → embed → store. Query time, every question: embed question → retrieve top-k → augment prompt → generate.
Formulas
| Dot product | \(\mathbf{a}\cdot\mathbf{b}=\sum_i a_i b_i\) |
| Length (L2 norm) | \(\lVert\mathbf{a}\rVert=\sqrt{\sum_i a_i^2}\) |
| Cosine similarity | \(\cos\theta=\dfrac{\mathbf{a}\cdot\mathbf{b}}{\lVert\mathbf{a}\rVert\lVert\mathbf{b}\rVert}\in[-1,1]\) |
| Normalise | \(\hat{\mathbf{a}}=\mathbf{a}/\lVert\mathbf{a}\rVert\), then \(\hat{\mathbf{a}}\cdot\hat{\mathbf{b}}=\cos\theta\) |
| L2 vs cosine (unit vectors) | \(\lVert\hat{\mathbf{a}}-\hat{\mathbf{b}}\rVert^2=2-2\cos\theta\) (same ranking) |
| Chroma cosine distance | \(d = 1-\cos\theta\) (smaller is better) |
| Hit rate@k | share of questions with ≥1 relevant doc in top k |
| Recall@k | mean over questions of (relevant docs found in top k) / (relevant docs) |
Core code (same names as the tutorial)
chunks = build_chunks(CORPUS, chunk_size=50, overlap=10) # step = chunk_size - overlap
embedder = Embedder("auto") # MiniLM (384-d), falls back to TF-IDF
index = build_index(chunks, embedder) # (n_chunks, dim), unit-length rows
q = embedder.encode([question])[0] # same embedder for queries!
scores = index @ q # cosine for every chunk
top = np.argsort(-scores)[:k] # best first
prompt = build_prompt(question, [chunks[i] for i in top])
metrics = evaluate_retrieval(EVAL_SET, index, chunks, embedder, k=3)
Defaults used in the course: word chunks of 50 with overlap 10 · all-MiniLM-L6-v2 (384 dimensions, input truncated after 256 word pieces) · k = 3 · temperature 0 for factual answers.
Measured on the Kestrelwood set (hit@1): MiniLM 50/10 = 1.00 · TF-IDF 50/10 = 0.90 · MiniLM 12/0 = 0.90 · precision@k fell from 1.00 (k=1) to 0.18 (k=10) while the prompt grew from 66 to 467 words.
Grounding prompt, three jobs: answer only from context · cite [DOC-ID] · fixed refusal sentence when the answer is missing.
Bedrock Converse (billed per token): client.converse(modelId=..., system=[{"text": ...}], messages=[{"role": "user", "content": [{"text": prompt}]}], inferenceConfig={"maxTokens": 300, "temperature": 0.0}) → answer at response["output"]["message"]["content"][0]["text"].
Debugging order: 1) Print the retrieved chunks and the exact prompt. 2) Right chunk missing? Then it's a retrieval problem: chunk size/overlap, embedding model, k, metadata filter, document not indexed. 3) Right chunk present but wrong answer? Then it's a generation problem: prompt instructions, too many chunks, the model. 4) Add the question to the evaluation set.
Mistakes to avoid: tiny chunks · no overlap · huge chunks · raw dot product on unnormalised vectors · different embedding models for index and query · large k "just in case" · trusting a similarity threshold · no evaluation set · re-indexing without re-running the evaluation (one new document can break old questions).
Part C: Quiz (15 questions)#
Suggested time: 25 minutes. Questions 1–8 are multiple choice (1 mark each), 9–14 are short answer (2 marks each), 15 is predict the output (2 marks). Total: 22 marks.
1. Which stage runs for every user question?A) Chunking the documentsB) Embedding the documentsC) Retrieving the top-k chunksD) Building the index
2. How many numbers does all-MiniLM-L6-v2 produce for one input text?A) 128B) 256C) 384D) 768
3. scores is a NumPy array of similarity scores. What does np.argsort(-scores)[:3] return?A) The three highest scoresB) The positions (indices) of the three highest scoresC) The positions of the three lowest scoresD) The three chunk texts
4. Two unit-length vectors have a dot product of 0.8. What is their cosine similarity?A) 0.2B) 0.64C) 0.8D) Cannot be determined
5. A Chroma collection configured with "space": "cosine" returns a distance of 0.25. What is the cosine similarity?A) 0.25B) 0.75C) 0.9375D) 4.0
6. Which change is most likely to reduce the input-token cost of each LLM call?A) Raise k from 3 to 10B) Lower k from 10 to 3C) Raise overlap from 10 to 20 wordsD) Switch from TF-IDF to MiniLM
7. In lab Checkpoint 6, the two retrieved HR-02 chunks repeat the phrase "a medical certificate from a registered doctor to the PeoplePortal". Why?A) A bug in searchB) Overlap between consecutive chunksC) HR-02 appears twice in the corpusD) The prompt template duplicates text
8. In the Bedrock Converse API, where do you put the instruction "Answer ONLY from the context provided"?A) modelIdB) systemC) inferenceConfigD) additionalModelResponseFieldPaths
9. (Short answer) In two or three sentences, explain why RAG reduces hallucination but does not eliminate it.
10. (Short answer) Compute the cosine similarity of \(\mathbf{a}=(1,0,1)\) and \(\mathbf{b}=(1,1,0)\) by hand. Show the dot product and both lengths.
11. (Short answer) A question has 3 relevant documents. The top 5 retrieved chunks come from 2 of them. What are hit@5 and recall@5 for this question?
12. (Short answer) Why must documents and queries be embedded by the same model? What error did the lab show when they weren't?
13. (Short answer) With TF-IDF, "How much time off do new fathers get?" retrieved the laptop policy instead of the parental leave policy. Explain why, and name the retrieval approach that avoided the mistake.
14. (Short answer) After a new "Public and floating holidays" document was added, "How many vacation days do I get each year?" started retrieving it instead of the annual leave policy. Why could this happen, and how would a team notice before users do?
15. (Predict the output) Using the lab's chunk_text, what does this print?
print(chunk_text("one two three four five six seven", chunk_size=3, overlap=1))
Answer key#
Show the answer key
| Q | Answer | Explanation |
|---|---|---|
| 1 | C | Chunking, document embedding and index building happen when documents change. Retrieval (plus embedding the question, augmenting and generating) happens per question. |
| 2 | C | 384 dimensions. 256 is the input limit in word pieces, a common distractor. |
| 3 | B | argsort returns indices, not values. Negating the scores turns ascending order into best-first. |
| 4 | C | For unit vectors the cosine denominator is \(1\times1\), so cosine equals the dot product. |
| 5 | B | Chroma's cosine distance is \(1-\cos\theta\), so \(\cos\theta = 1-0.25 = 0.75\). |
| 6 | B | Fewer retrieved chunks means fewer input tokens per call. More overlap adds text; the embedding choice doesn't change prompt length. |
| 7 | B | With overlap 10, each chunk repeats the last 10 words of the previous chunk. |
| 8 | B | system takes a list of content blocks for instructions. inferenceConfig holds maxTokens, temperature, topP and stopSequences. |
| 9 | Model answer | RAG puts relevant passages from your documents into the prompt, so the model can answer from facts instead of guessing (1 mark). It can still fail: retrieval may return the wrong or incomplete chunk, and the model may ignore, misread or go beyond the context (1 mark). |
| 10 | 0.5 | \(\mathbf{a}\cdot\mathbf{b} = 1\cdot1 + 0\cdot1 + 1\cdot0 = 1\); \(\lVert\mathbf{a}\rVert = \lVert\mathbf{b}\rVert = \sqrt{2}\); \(\cos\theta = 1/(\sqrt2\cdot\sqrt2) = 1/2 = 0.5\) (1 mark for working, 1 for the answer). |
| 11 | hit@5 = 1, recall@5 = 2/3 ≈ 0.67 | At least one relevant document was found, so hit = 1; 2 of 3 were found, so recall = 0.67. |
| 12 | Model answer | Each model defines its own vector space (and often its own dimension), so distances between vectors from different models are meaningless (1 mark). The lab raised ValueError: matmul: ... (size 525 is different from 384) when a TF-IDF query was searched against a MiniLM index (1 mark). |
| 13 | Model answer | TF-IDF matches words, not meaning: after stop words are removed, "new" matched "Every new joiner" in the laptop policy and outscored the single match on "fathers" (0.087 against 0.070) (1 mark). Dense embeddings (MiniLM with cosine similarity) ranked the parental leave policy first (1 mark). |
| 14 | Model answer | Embeddings place "vacation" close to "holidays", so the new document scored higher than the annual leave policy for that question (1 mark). A labelled evaluation set re-run after every corpus change catches it: hit@1 fell from 1.00 to 0.91 in the tutorial's run (1 mark). |
| 15 | ['one two three', 'three four five', 'five six seven'] |
step = 3 - 1 = 2, so chunks start at words 0, 2 and 4. The third chunk reaches the last word (4 + 3 = 7 words), so the loop stops. Output verified by running the code. |
Grading guide: 18–22 ready for the vector database module · 12–17 revisit the cheat sheet and redo Step 7 · below 12 repeat the lab with the tutorial open alongside.
Continue
Read the full walkthrough: Build a RAG Application in Python, Step by Step
Learn it live: RAG & LLM Applications training · WhatsApp +91 70492 35525