What you will learn#
By the end you will be able to:
- Explain why a large language model (LLM) on its own hallucinates about private or recent information, and how RAG addresses it.
- Describe the seven stages of a RAG pipeline: ingest, chunk, embed, store, retrieve, augment, generate.
- Compute cosine similarity by hand and explain why normalised vectors make it a plain dot product.
- Write a chunker with overlap, build an in-memory vector index and search it with NumPy.
- Assemble a grounded prompt and send it to an LLM (Amazon Bedrock Converse API shown, any chat model works).
- Move the same index to FAISS or Chroma.
- Measure retrieval with hit rate@k and recall@k on a hand-labelled question set, and use the numbers to choose chunk size and top-k.
Prerequisites#
- Python 3.10 or newer recommended. I ran everything on Python 3.13.5.
- Basic Python: functions, lists, dictionaries, f-strings,
pip install. - NumPy basics are helpful but not required; every array operation is explained.
- No GPU, no cloud account, no API key for the retrieval part.
- Optional, for the generation step only: an AWS account with Amazon Bedrock access and configured credentials. This part costs money.
- New to embeddings and LLMs? Read section 5 of my AI & Machine Learning Foundations guide first. It covers tokens, embeddings and the RAG mental model in plain language.
Setup#
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# Optional on Linux: install the CPU-only PyTorch wheel first to avoid a large CUDA download
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install numpy scikit-learn sentence-transformers # retrieval (required)
pip install faiss-cpu chromadb # vector database step (optional)
pip install boto3 # Amazon Bedrock generation step (optional)
Versions used for every output in this tutorial:
| Package | Version |
|---|---|
| Python | 3.13.5 |
| numpy | 2.5.3 |
| scikit-learn | 1.9.1 |
| sentence-transformers | 6.1.0 |
| torch | 2.14.0+cpu |
| transformers | 5.17.0 |
| faiss-cpu | 1.15.1 |
| chromadb | 1.5.9 |
| boto3 | 1.43.102 |
| botocore | 1.43.102 |
The first run downloads the all-MiniLM-L6-v2 embedding model from Hugging Face. If your network blocks that download, the code automatically falls back to scikit-learn TF-IDF vectors and tells you so (you'll see that fallback output later).
Why an LLM alone is not enough#
An LLM generates text by predicting likely next tokens from patterns learned during training. That design has three consequences for business questions:
- It has never seen your private data. Your leave policy, runbooks and contracts were not in the training set, so the model cannot know them.
- Its knowledge stops at a training cut-off. A policy that changed last month is invisible to it.
- It answers anyway. Nothing in next-token prediction forces the model to say "I don't know". When the facts are missing it produces fluent text that fits the pattern of an answer. This is what people call hallucination.
Fine-tuning can change a model's style and behaviour, but it is a poor way to store facts that change often: every policy update would mean another training run, and you still couldn't point to the source of an answer.
The core idea of RAG#
RAG keeps the model's weights unchanged and gives it an open-book exam instead. At question time you:
- Retrieve the few passages from your documents that are most relevant to the question.
- Augment the prompt by pasting those passages in as context, with an instruction to answer only from them.
- Generate the answer with the LLM, which now reads your facts instead of guessing.
The term comes from Lewis et al. (2020), Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, which combined a neural retriever over a dense vector index of Wikipedia with a sequence-to-sequence generator. They called the index the model's non-parametric memory and showed it could be swapped to update the model's knowledge without retraining. Modern RAG applications keep that idea but usually use a general-purpose chat LLM and don't train the retriever and generator together.
What RAG gives you:
- Current answers: update a document, re-index it, and the next answer uses the new text.
- Traceability: every answer can cite the chunk it came from.
- Access control: you can filter what gets retrieved per user before the model sees anything.
What RAG does not give you: a guarantee. If retrieval returns the wrong chunk, the model answers from the wrong chunk. That is why this tutorial spends as much time on measuring retrieval as on building it.
How a RAG pipeline works#
| Stage | What happens | In our code |
|---|---|---|
| Ingest | Load text from PDFs, wikis, databases. Keep metadata (id, title, date, owner). | CORPUS list of dicts |
| Chunk | Split long documents into passages small enough to be specific. | chunk_text(), build_chunks() |
| Embed | Turn each chunk into a vector of numbers that captures its meaning. | Embedder.encode() |
| Store | Keep vectors and chunk text together so you can search them. | build_index() returns a NumPy matrix |
| Retrieve | Embed the question and find the most similar chunk vectors. | search() |
| Augment | Put the retrieved chunks and the question into one prompt. | build_prompt() |
| Generate | Send the prompt to an LLM and return its answer. | generate_with_bedrock() or any llm_fn |
The first four stages run offline whenever documents change. The last three run for every question, so they need to be fast.
The maths: embeddings, dot product and cosine similarity#
Embeddings are vectors#
An embedding model maps a piece of text to a fixed-length list of numbers, a vector. The model we use, all-MiniLM-L6-v2, outputs 384 numbers per text. It was trained so that texts with similar meaning get vectors that point in similar directions. "I forgot my password" and "How do I reset my login?" share almost no words, but in my run their cosine similarity (explained below) was 0.62, against 0.08 between "I forgot my password" and a sentence about hotel limits.
So "find passages relevant to this question" becomes a geometry problem: find the chunk vectors that point in the most similar direction to the question vector. We need a number that measures "similar direction".
Dot product#
- \(\mathbf{a}\) and \(\mathbf{b}\) are two vectors, for example a question embedding and a chunk embedding.
- \(n\) is the number of dimensions (384 for MiniLM).
- \(a_i\) and \(b_i\) are the \(i\)-th numbers in each vector.
- \(\sum_{i=1}^{n}\) means "add up the terms for \(i = 1\) to \(n\)". Multiply matching positions, then add.
The dot product is large when the vectors point the same way, but it also grows with the vectors' length, which has nothing to do with meaning.
Vector length (the L2 norm)#
- \(\lVert\mathbf{a}\rVert\) is the length (Euclidean or L2 norm) of \(\mathbf{a}\): square every component, add them up, take the square root. It is Pythagoras extended to \(n\) dimensions.
Cosine similarity#
- \(\theta\) (theta) is the angle between the two vectors.
- The numerator \(\mathbf{a}\cdot\mathbf{b}\) is the dot product defined above.
- The denominator \(\lVert\mathbf{a}\rVert\,\lVert\mathbf{b}\rVert\) is the product of the two lengths. Dividing by it removes the effect of length and leaves only direction.
- The result lies between \(-1\) and \(1\): \(1\) means the same direction, \(0\) means unrelated (perpendicular), \(-1\) means opposite.
A worked example by hand#
Take a toy 3-dimensional "question" vector and two "chunk" vectors:
- \(\mathbf{q}\) is the question vector; \(\mathbf{d}_1\) and \(\mathbf{d}_2\) are two candidate chunk vectors.
Dot products:
By dot product, \(\mathbf{d}_2\) wins (12 > 8).
Lengths:
Cosine similarities:
By cosine, \(\mathbf{d}_1\) wins. The angle between \(\mathbf{q}\) and \(\mathbf{d}_1\) is about 27°, while \(\mathbf{d}_2\) is about 48° away. \(\mathbf{d}_2\) won the dot-product contest only because it is twice as long. With raw count-style vectors, a long chunk that repeats one word behaves like \(\mathbf{d}_2\): a large dot product without being more relevant.
Normalise once, then the dot product is the cosine#
Divide every vector by its length:
- \(\hat{\mathbf{a}}\) ("a-hat") is the unit vector pointing the same way as \(\mathbf{a}\), with length exactly 1.
For unit vectors the denominator of the cosine formula is \(1 \times 1\), so
Check with our numbers: \(\hat{\mathbf{q}}=(\tfrac13,\tfrac23,\tfrac23)\) and \(\hat{\mathbf{d}}_1=(\tfrac23,\tfrac13,\tfrac23)\), so \(\hat{\mathbf{q}}\cdot\hat{\mathbf{d}}_1=\tfrac{2+2+4}{9}=\tfrac89\). Same answer.
This is why production systems normalise embeddings at indexing time: search then needs only multiplications and additions. It also explains a useful identity for unit vectors:
- The left side is the squared Euclidean (L2) distance between the two unit vectors.
Smaller distance means larger cosine, so for normalised vectors, ranking by L2 distance and ranking by cosine give the same order.
Searching every chunk at once#
Stack the \(N\) unit-length chunk vectors as the rows of a matrix \(E\) (shape \(N \times d\)). Then
- \(E\) is the index matrix, one row per chunk; \(N\) is the number of chunks and \(d\) the embedding dimension.
- \(\hat{\mathbf{q}}\) is the normalised question vector (length \(d\)).
- \(\mathbf{s}\) is a vector of \(N\) scores; \(s_j\) is the cosine similarity between the question and chunk \(j\).
Retrieval is "compute \(\mathbf{s}\), return the \(k\) largest". In NumPy that is index @ q followed by np.argsort.
Build it from scratch#
The complete script is rag_from_scratch.py: download the full code bundle (rag-tutorial-code.zip) or the main script only (rag_from_scratch.py), or read it in full below. The sections below walk through it in order.
Read the complete rag_from_scratch.py (400 lines)
"""
rag_from_scratch.py
Build a Retrieval-Augmented Generation (RAG) pipeline in plain Python + NumPy.
Companion code for "Build a RAG Application in Python, Step by Step" (pushpjeet.com).
What runs without any API key or cloud account:
chunking -> embeddings -> in-memory vector index -> cosine-similarity search
-> prompt assembly -> retrieval evaluation (hit rate, recall@k)
What does NOT run by default:
generate_with_bedrock() calls Amazon Bedrock (a paid AWS service). It is only
called if you pass --generate and supply your own model ID. It needs AWS
credentials and incurs charges.
The corpus below is for a FICTIONAL company ("Kestrelwood Systems").
All names, numbers, e-mail addresses and URLs are made up for teaching.
Usage:
python rag_from_scratch.py # sentence-transformers (falls back to TF-IDF)
python rag_from_scratch.py --backend tfidf # force the TF-IDF fallback
python rag_from_scratch.py --generate --model-id <YOUR_MODEL_ID> --region <REGION>
"""
from __future__ import annotations
import argparse
import numpy as np
# ---------------------------------------------------------------------------
# 1. Sample corpus (fictional)
# ---------------------------------------------------------------------------
CORPUS = [
{
"doc_id": "HR-01",
"title": "Annual leave policy",
"text": (
"Every full-time employee at Kestrelwood Systems receives 24 days of paid annual "
"leave per calendar year. Leave accrues at 2 days per month. You can carry forward "
"a maximum of 10 unused days into the next calendar year; any balance above 10 days "
"lapses on 31 December. Apply for leave in the PeoplePortal at least 5 working days "
"in advance. Your reporting manager approves or rejects the request within 2 working days."
),
},
{
"doc_id": "HR-02",
"title": "Sick leave policy",
"text": (
"Employees receive 12 days of paid sick leave per calendar year. Sick leave cannot be "
"carried forward. Inform your manager before 10:00 AM on the day you are unwell. "
"If you are absent for more than 2 consecutive days, upload a medical certificate "
"from a registered doctor to the PeoplePortal within 3 days of returning to work."
),
},
{
"doc_id": "HR-03",
"title": "Hybrid and remote work",
"text": (
"Kestrelwood follows a hybrid model. You may work remotely up to 3 days per week. "
"Wednesday is the team anchor day and everyone works from the office. Core "
"collaboration hours are 11:00 to 16:00 IST, when you must be reachable on chat. "
"When working outside the office, always connect through the company VPN."
),
},
{
"doc_id": "FIN-01",
"title": "Expense reimbursement",
"text": (
"Submit expense claims in ExpenseDesk within 30 days of spending the money. Attach "
"an itemised receipt for every line item; card statements are not accepted as receipts. "
"Approved claims are paid with the next monthly payroll. Any single claim above "
"INR 25,000 needs approval from your department head in addition to your manager."
),
},
{
"doc_id": "FIN-02",
"title": "Business travel policy",
"text": (
"Book all business travel through the internal travel desk at least 7 days before the "
"trip. Flights shorter than 6 hours must be booked in economy class. The hotel limit "
"is INR 6,000 per night in metro cities and INR 4,000 per night elsewhere. A daily "
"allowance of INR 1,500 covers meals and local transport."
),
},
{
"doc_id": "IT-01",
"title": "Passwords and account security",
"text": (
"Passwords must be at least 14 characters long. Multi-factor authentication is "
"mandatory for every company account. If you forget your password, reset it yourself "
"at the self-service portal id.kestrelwood.example using your registered phone. "
"Never share one-time passcodes with anyone. The IT team will never ask for your password."
),
},
{
"doc_id": "IT-02",
"title": "Laptops and devices",
"text": (
"Every new joiner receives a company laptop on day one. All laptops use full-disk "
"encryption and are managed centrally. If your laptop or phone is lost or stolen, "
"report it to the IT helpdesk on extension 4040 within 2 hours so the device can be "
"locked and wiped remotely. Personal devices may access email only through the managed browser."
),
},
{
"doc_id": "LND-01",
"title": "Learning and certification budget",
"text": (
"Each employee has a learning budget of INR 50,000 per financial year for courses, "
"books and professional certification exams. Get written pre-approval from your manager "
"before you pay. Claim the cost through ExpenseDesk with the invoice and proof of completion. "
"If you leave the company within 12 months of a reimbursement, you repay 50 percent of it."
),
},
{
"doc_id": "HR-04",
"title": "Parental leave",
"text": (
"Birth mothers receive 26 weeks of paid maternity leave. Fathers, partners and "
"adoptive parents receive 4 weeks of paid parental leave, which must be taken within "
"6 months of the birth or adoption. After parental leave, employees can request a "
"phased return with reduced hours for up to 8 weeks."
),
},
{
"doc_id": "SEC-01",
"title": "Data classification and AI tool usage",
"text": (
"Company information is classified as Public, Internal, Confidential or Restricted. "
"Customer personal data is always Restricted. Never paste Confidential or Restricted "
"data into external AI chatbots or public websites. Use only the approved internal "
"assistant for work involving customer data, and report any accidental disclosure to "
"security@kestrelwood.example immediately."
),
},
]
# Hand-labelled evaluation set: question -> set of doc_ids that contain the answer.
EVAL_SET = [
{"question": "How many vacation days do I get each year?", "relevant": {"HR-01"}},
{"question": "Can I carry unused leave into next year?", "relevant": {"HR-01"}},
{"question": "Do I need a doctor's note if I am ill for three days?", "relevant": {"HR-02"}},
{"question": "I forgot my password. How do I reset it?", "relevant": {"IT-01"}},
{"question": "My laptop was stolen at the airport. What should I do?", "relevant": {"IT-02"}},
{"question": "What is the hotel limit when I travel to Mumbai?", "relevant": {"FIN-02"}},
{"question": "Can I use ChatGPT to summarise a customer's complaint?", "relevant": {"SEC-01"}},
{"question": "How much time off do new fathers get?", "relevant": {"HR-04"}},
{"question": "Will the company pay for my AWS certification exam?", "relevant": {"LND-01"}},
{
"question": "How do I book flights for a client visit and claim the costs afterwards?",
"relevant": {"FIN-02", "FIN-01"},
},
]
# ---------------------------------------------------------------------------
# 2. Chunking
# ---------------------------------------------------------------------------
def chunk_text(text: str, chunk_size: int = 50, overlap: int = 10) -> list[str]:
"""Split text into chunks of `chunk_size` words; consecutive chunks share `overlap` words."""
if chunk_size <= 0:
raise ValueError("chunk_size must be positive")
if not 0 <= overlap < chunk_size:
raise ValueError("overlap must be >= 0 and smaller than chunk_size")
words = text.split()
step = chunk_size - overlap
chunks = []
for start in range(0, len(words), step):
chunks.append(" ".join(words[start:start + chunk_size]))
if start + chunk_size >= len(words):
break
return chunks
def build_chunks(corpus: list[dict], chunk_size: int = 50, overlap: int = 10) -> list[dict]:
"""Chunk every document and keep metadata (doc_id, title) with each chunk."""
chunks = []
for doc in corpus:
for i, piece in enumerate(chunk_text(doc["text"], chunk_size, overlap)):
chunks.append({
"chunk_id": f"{doc['doc_id']}#{i}",
"doc_id": doc["doc_id"],
"title": doc["title"],
"text": piece,
})
return chunks
# ---------------------------------------------------------------------------
# 3. Embeddings
# ---------------------------------------------------------------------------
class Embedder:
"""Turns text into L2-normalised vectors.
backend="auto" -> try sentence-transformers all-MiniLM-L6-v2, fall back to TF-IDF
backend="minilm" -> sentence-transformers only
backend="tfidf" -> scikit-learn TF-IDF only (no download needed)
"""
MODEL_NAME = "sentence-transformers/all-MiniLM-L6-v2"
def __init__(self, backend: str = "auto"):
self.backend = None
self.model = None
if backend in ("auto", "minilm"):
try:
from sentence_transformers import SentenceTransformer
self.model = SentenceTransformer(self.MODEL_NAME)
self.backend = "minilm"
except Exception as exc: # no package, no network, blocked download...
if backend == "minilm":
raise
print(f"[info] sentence-transformers unavailable ({exc!r}); using TF-IDF.")
if self.backend is None:
from sklearn.feature_extraction.text import TfidfVectorizer
self.model = TfidfVectorizer(stop_words="english", ngram_range=(1, 2))
self.backend = "tfidf"
self._fitted = self.backend == "minilm"
def fit(self, texts: list[str]) -> "Embedder":
"""TF-IDF must learn its vocabulary from the corpus; MiniLM is already trained."""
if self.backend == "tfidf":
self.model.fit(texts)
self._fitted = True
return self
def encode(self, texts: list[str]) -> np.ndarray:
if not self._fitted:
raise RuntimeError("Call fit() on the corpus before encode() when using TF-IDF.")
if self.backend == "minilm":
vectors = self.model.encode(texts, convert_to_numpy=True)
else:
vectors = self.model.transform(texts).toarray()
vectors = vectors.astype(np.float32)
norms = np.linalg.norm(vectors, axis=1, keepdims=True)
norms[norms == 0] = 1.0 # avoid division by zero for empty vectors
return vectors / norms # unit length -> dot product == cosine similarity
# ---------------------------------------------------------------------------
# 4. Vector index and similarity search
# ---------------------------------------------------------------------------
def cosine_similarity(a: np.ndarray, b: np.ndarray) -> float:
"""cos(theta) = (a . b) / (||a|| * ||b||)"""
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
def build_index(chunks: list[dict], embedder: Embedder) -> np.ndarray:
"""Embed every chunk. Row i of the returned matrix is the vector for chunks[i]."""
texts = [c["text"] for c in chunks]
embedder.fit(texts)
return embedder.encode(texts)
def search(query: str, index: np.ndarray, chunks: list[dict], embedder: Embedder, k: int = 3) -> list[dict]:
"""Return the k chunks most similar to the query, best first."""
q = embedder.encode([query])[0]
scores = index @ q # one dot product per chunk (all vectors are unit length)
top = np.argsort(-scores)[:k] # indices of the k highest scores
return [{**chunks[i], "score": float(scores[i])} for i in top]
# ---------------------------------------------------------------------------
# 5. Prompt assembly
# ---------------------------------------------------------------------------
SYSTEM_PROMPT = (
"You are the HR and IT help assistant for Kestrelwood Systems. "
"Answer ONLY from the context provided. Cite the source id in square brackets, e.g. [HR-01]. "
"If the context does not contain the answer, reply exactly: "
"\"I don't know based on the company documents.\""
)
def build_prompt(question: str, retrieved: list[dict]) -> str:
"""Put the retrieved chunks and the question into one user message."""
context = "\n\n".join(f"[{r['doc_id']}] {r['title']}\n{r['text']}" for r in retrieved)
return f"Context:\n{context}\n\nQuestion: {question}\nAnswer:"
# ---------------------------------------------------------------------------
# 6. Generation (pluggable). NOT executed by default: Amazon Bedrock is a paid service.
# ---------------------------------------------------------------------------
def generate_with_bedrock(prompt: str, system_prompt: str = SYSTEM_PROMPT,
model_id: str = "<YOUR_BEDROCK_MODEL_ID>",
region: str = "us-east-1", client=None) -> str:
"""Send the RAG prompt to a Bedrock model with the Converse API and return the answer text.
COSTS MONEY: each call is billed per input and output token. Requires AWS credentials
with permission for bedrock:InvokeModel. Replace model_id with a model ID (or inference
profile ID) that is enabled in your account and region. Pass `client` to reuse a
boto3 client (or a stubbed one in tests).
"""
if client is None:
import boto3
client = boto3.client("bedrock-runtime", region_name=region)
response = client.converse(
modelId=model_id,
system=[{"text": system_prompt}],
messages=[{"role": "user", "content": [{"text": prompt}]}],
inferenceConfig={"maxTokens": 300, "temperature": 0.0},
)
# The reply is a list of content blocks; keep the text ones (some models add reasoning blocks).
blocks = response["output"]["message"]["content"]
return "".join(block["text"] for block in blocks if "text" in block)
def rag_answer(question: str, index: np.ndarray, chunks: list[dict], embedder: Embedder,
k: int = 3, llm_fn=None) -> dict:
"""Retrieve -> augment -> (optionally) generate.
llm_fn is any function that takes (prompt, system_prompt) and returns text,
so you can swap Bedrock for any other chat LLM without touching retrieval.
"""
retrieved = search(question, index, chunks, embedder, k=k)
prompt = build_prompt(question, retrieved)
answer = llm_fn(prompt, SYSTEM_PROMPT) if llm_fn else None
return {"question": question, "retrieved": retrieved, "prompt": prompt, "answer": answer}
# ---------------------------------------------------------------------------
# 7. Retrieval evaluation
# ---------------------------------------------------------------------------
def evaluate_retrieval(eval_set: list[dict], index: np.ndarray, chunks: list[dict],
embedder: Embedder, k: int = 3, verbose: bool = False) -> dict:
"""Hit rate@k: share of questions with at least one relevant doc in the top k.
Recall@k: average share of each question's relevant docs found in the top k."""
hits, recalls = [], []
for item in eval_set:
results = search(item["question"], index, chunks, embedder, k=k)
found = {r["doc_id"] for r in results} & item["relevant"]
hits.append(1.0 if found else 0.0)
recalls.append(len(found) / len(item["relevant"]))
if verbose:
got = ", ".join(r["doc_id"] for r in results)
mark = "HIT " if found else "MISS"
print(f" {mark} want={sorted(item['relevant'])} got=[{got}] {item['question']}")
return {"k": k, "hit_rate": float(np.mean(hits)), "recall": float(np.mean(recalls))}
# ---------------------------------------------------------------------------
# Demo
# ---------------------------------------------------------------------------
def main() -> None:
parser = argparse.ArgumentParser(description="RAG from scratch demo")
parser.add_argument("--backend", default="auto", choices=["auto", "minilm", "tfidf"])
parser.add_argument("--generate", action="store_true", help="call Amazon Bedrock (costs money)")
parser.add_argument("--model-id", default="<YOUR_BEDROCK_MODEL_ID>")
parser.add_argument("--region", default="us-east-1")
args = parser.parse_args()
print("== Step 0: cosine similarity by hand ==")
q, d1, d2 = np.array([1, 2, 2]), np.array([2, 1, 2]), np.array([0, 6, 0])
print(f"dot(q,d1)={np.dot(q, d1)} dot(q,d2)={np.dot(q, d2)}")
print(f"cos(q,d1)={cosine_similarity(q, d1):.3f} cos(q,d2)={cosine_similarity(q, d2):.3f}")
print("\n== Step 1: chunking ==")
chunks = build_chunks(CORPUS, chunk_size=50, overlap=10)
print(f"{len(CORPUS)} documents -> {len(chunks)} chunks (chunk_size=50 words, overlap=10)")
for c in chunks[:2]:
print(f" {c['chunk_id']}: {c['text'][:90]}...")
print("\n== Step 2: embeddings + index ==")
embedder = Embedder(args.backend)
index = build_index(chunks, embedder)
print(f"backend={embedder.backend} index shape={index.shape} "
f"first row norm={np.linalg.norm(index[0]):.3f}")
print("\n== Step 3: search ==")
for question in ["I forgot my password. How do I reset it?",
"Can I use ChatGPT to summarise a customer's complaint?"]:
print(f"Q: {question}")
for r in search(question, index, chunks, embedder, k=3):
print(f" {r['score']:.3f} {r['chunk_id']:<9} {r['title']}")
print("\n== Step 4: prompt assembly ==")
result = rag_answer("How many vacation days do I get each year?", index, chunks, embedder, k=2)
print(result["prompt"])
print("\n== Step 5: a question the documents cannot answer ==")
unanswerable = "What is Kestrelwood's share price today?"
print(f"Q: {unanswerable}")
for r in search(unanswerable, index, chunks, embedder, k=3):
print(f" {r['score']:.3f} {r['chunk_id']:<9} {r['title']}")
print("\n== Step 6: retrieval evaluation ==")
for k in (1, 3):
m = evaluate_retrieval(EVAL_SET, index, chunks, embedder, k=k, verbose=(k == 1))
print(f"k={m['k']}: hit_rate={m['hit_rate']:.2f} recall={m['recall']:.2f}")
if args.generate:
print("\n== Step 7: generation with Amazon Bedrock (billed) ==")
llm = lambda p, s: generate_with_bedrock(p, s, model_id=args.model_id, region=args.region)
out = rag_answer("How many vacation days do I get each year?", index, chunks, embedder,
k=3, llm_fn=llm)
print(out["answer"])
if __name__ == "__main__":
main()
First, the data.
A small corpus#
We use ten short HR and IT policy documents for Kestrelwood Systems, a fictional company. Every name, number, email address and URL in them is made up. Each document has an id, which we will use for citations and evaluation.
Show the full corpus and the labelled evaluation questions (copy into your script)
CORPUS = [
{
"doc_id": "HR-01",
"title": "Annual leave policy",
"text": (
"Every full-time employee at Kestrelwood Systems receives 24 days of paid annual "
"leave per calendar year. Leave accrues at 2 days per month. You can carry forward "
"a maximum of 10 unused days into the next calendar year; any balance above 10 days "
"lapses on 31 December. Apply for leave in the PeoplePortal at least 5 working days "
"in advance. Your reporting manager approves or rejects the request within 2 working days."
),
},
{
"doc_id": "HR-02",
"title": "Sick leave policy",
"text": (
"Employees receive 12 days of paid sick leave per calendar year. Sick leave cannot be "
"carried forward. Inform your manager before 10:00 AM on the day you are unwell. "
"If you are absent for more than 2 consecutive days, upload a medical certificate "
"from a registered doctor to the PeoplePortal within 3 days of returning to work."
),
},
{
"doc_id": "HR-03",
"title": "Hybrid and remote work",
"text": (
"Kestrelwood follows a hybrid model. You may work remotely up to 3 days per week. "
"Wednesday is the team anchor day and everyone works from the office. Core "
"collaboration hours are 11:00 to 16:00 IST, when you must be reachable on chat. "
"When working outside the office, always connect through the company VPN."
),
},
{
"doc_id": "FIN-01",
"title": "Expense reimbursement",
"text": (
"Submit expense claims in ExpenseDesk within 30 days of spending the money. Attach "
"an itemised receipt for every line item; card statements are not accepted as receipts. "
"Approved claims are paid with the next monthly payroll. Any single claim above "
"INR 25,000 needs approval from your department head in addition to your manager."
),
},
{
"doc_id": "FIN-02",
"title": "Business travel policy",
"text": (
"Book all business travel through the internal travel desk at least 7 days before the "
"trip. Flights shorter than 6 hours must be booked in economy class. The hotel limit "
"is INR 6,000 per night in metro cities and INR 4,000 per night elsewhere. A daily "
"allowance of INR 1,500 covers meals and local transport."
),
},
{
"doc_id": "IT-01",
"title": "Passwords and account security",
"text": (
"Passwords must be at least 14 characters long. Multi-factor authentication is "
"mandatory for every company account. If you forget your password, reset it yourself "
"at the self-service portal id.kestrelwood.example using your registered phone. "
"Never share one-time passcodes with anyone. The IT team will never ask for your password."
),
},
{
"doc_id": "IT-02",
"title": "Laptops and devices",
"text": (
"Every new joiner receives a company laptop on day one. All laptops use full-disk "
"encryption and are managed centrally. If your laptop or phone is lost or stolen, "
"report it to the IT helpdesk on extension 4040 within 2 hours so the device can be "
"locked and wiped remotely. Personal devices may access email only through the managed browser."
),
},
{
"doc_id": "LND-01",
"title": "Learning and certification budget",
"text": (
"Each employee has a learning budget of INR 50,000 per financial year for courses, "
"books and professional certification exams. Get written pre-approval from your manager "
"before you pay. Claim the cost through ExpenseDesk with the invoice and proof of completion. "
"If you leave the company within 12 months of a reimbursement, you repay 50 percent of it."
),
},
{
"doc_id": "HR-04",
"title": "Parental leave",
"text": (
"Birth mothers receive 26 weeks of paid maternity leave. Fathers, partners and "
"adoptive parents receive 4 weeks of paid parental leave, which must be taken within "
"6 months of the birth or adoption. After parental leave, employees can request a "
"phased return with reduced hours for up to 8 weeks."
),
},
{
"doc_id": "SEC-01",
"title": "Data classification and AI tool usage",
"text": (
"Company information is classified as Public, Internal, Confidential or Restricted. "
"Customer personal data is always Restricted. Never paste Confidential or Restricted "
"data into external AI chatbots or public websites. Use only the approved internal "
"assistant for work involving customer data, and report any accidental disclosure to "
"security@kestrelwood.example immediately."
),
},
]
# Hand-labelled evaluation set: question -> set of doc_ids that contain the answer.
EVAL_SET = [
{"question": "How many vacation days do I get each year?", "relevant": {"HR-01"}},
{"question": "Can I carry unused leave into next year?", "relevant": {"HR-01"}},
{"question": "Do I need a doctor's note if I am ill for three days?", "relevant": {"HR-02"}},
{"question": "I forgot my password. How do I reset it?", "relevant": {"IT-01"}},
{"question": "My laptop was stolen at the airport. What should I do?", "relevant": {"IT-02"}},
{"question": "What is the hotel limit when I travel to Mumbai?", "relevant": {"FIN-02"}},
{"question": "Can I use ChatGPT to summarise a customer's complaint?", "relevant": {"SEC-01"}},
{"question": "How much time off do new fathers get?", "relevant": {"HR-04"}},
{"question": "Will the company pay for my AWS certification exam?", "relevant": {"LND-01"}},
{
"question": "How do I book flights for a client visit and claim the costs afterwards?",
"relevant": {"FIN-02", "FIN-01"},
},
]
The EVAL_SET at the end is a list of questions I labelled by hand with the document(s) that contain the answer. We will use it in the evaluation section. Notice that most questions deliberately use different words from the documents ("vacation" versus "annual leave", "ChatGPT" versus "external AI chatbots", "Mumbai" versus "metro cities"). Real users do this all the time.
Step 1: Chunking#
Why chunk at all? An embedding compresses a whole text into one vector. Embed a 20-page policy as one vector and the details average out, so the vector is vaguely about everything and specifically about nothing. MiniLM also truncates input longer than 256 word pieces, so text beyond that point would not be represented at all. Small chunks keep each vector specific.
Why overlap? A hard cut can split one fact across two chunks ("...absent for more than | 2 consecutive days..."). Repeating a few words at each boundary gives each fact a better chance of landing whole in at least one chunk.
def chunk_text(text: str, chunk_size: int = 50, overlap: int = 10) -> list[str]:
"""Split text into chunks of `chunk_size` words; consecutive chunks share `overlap` words."""
if chunk_size <= 0:
raise ValueError("chunk_size must be positive")
if not 0 <= overlap < chunk_size:
raise ValueError("overlap must be >= 0 and smaller than chunk_size")
words = text.split()
step = chunk_size - overlap
chunks = []
for start in range(0, len(words), step):
chunks.append(" ".join(words[start:start + chunk_size]))
if start + chunk_size >= len(words):
break
return chunks
def build_chunks(corpus: list[dict], chunk_size: int = 50, overlap: int = 10) -> list[dict]:
"""Chunk every document and keep metadata (doc_id, title) with each chunk."""
chunks = []
for doc in corpus:
for i, piece in enumerate(chunk_text(doc["text"], chunk_size, overlap)):
chunks.append({
"chunk_id": f"{doc['doc_id']}#{i}",
"doc_id": doc["doc_id"],
"title": doc["title"],
"text": piece,
})
return chunks
How it works: step = chunk_size - overlap is how far the window moves each time. With chunk_size=50 and overlap=10, chunks start at words 0, 40, 80, and so on, and each chunk repeats the last 10 words of the previous one. build_chunks keeps doc_id and title with every chunk, because retrieval returns chunks but citations and evaluation need documents.
I split on words to keep the code readable. Production systems usually count tokens and prefer to cut at sentence or heading boundaries; the idea is the same.
Expected output (from python rag_from_scratch.py):
== Step 1: chunking ==
10 documents -> 17 chunks (chunk_size=50 words, overlap=10)
HR-01#0: Every full-time employee at Kestrelwood Systems receives 24 days of paid annual leave per ...
HR-01#1: above 10 days lapses on 31 December. Apply for leave in the PeoplePortal at least 5 workin...
Documents of 48 to 72 words become one or two chunks each, 17 in total.
Step 2: Embeddings#
class Embedder:
"""Turns text into L2-normalised vectors.
backend="auto" -> try sentence-transformers all-MiniLM-L6-v2, fall back to TF-IDF
backend="minilm" -> sentence-transformers only
backend="tfidf" -> scikit-learn TF-IDF only (no download needed)
"""
MODEL_NAME = "sentence-transformers/all-MiniLM-L6-v2"
def __init__(self, backend: str = "auto"):
self.backend = None
self.model = None
if backend in ("auto", "minilm"):
try:
from sentence_transformers import SentenceTransformer
self.model = SentenceTransformer(self.MODEL_NAME)
self.backend = "minilm"
except Exception as exc: # no package, no network, blocked download...
if backend == "minilm":
raise
print(f"[info] sentence-transformers unavailable ({exc!r}); using TF-IDF.")
if self.backend is None:
from sklearn.feature_extraction.text import TfidfVectorizer
self.model = TfidfVectorizer(stop_words="english", ngram_range=(1, 2))
self.backend = "tfidf"
self._fitted = self.backend == "minilm"
def fit(self, texts: list[str]) -> "Embedder":
"""TF-IDF must learn its vocabulary from the corpus; MiniLM is already trained."""
if self.backend == "tfidf":
self.model.fit(texts)
self._fitted = True
return self
def encode(self, texts: list[str]) -> np.ndarray:
if not self._fitted:
raise RuntimeError("Call fit() on the corpus before encode() when using TF-IDF.")
if self.backend == "minilm":
vectors = self.model.encode(texts, convert_to_numpy=True)
else:
vectors = self.model.transform(texts).toarray()
vectors = vectors.astype(np.float32)
norms = np.linalg.norm(vectors, axis=1, keepdims=True)
norms[norms == 0] = 1.0 # avoid division by zero for empty vectors
return vectors / norms # unit length -> dot product == cosine similarity
Points worth understanding:
all-MiniLM-L6-v2is a small sentence-transformers model that maps text to a 384-dimensional vector. It runs quickly on a CPU.- Fallback: if the package or the model download is unavailable, the class switches to TF-IDF from scikit-learn. TF-IDF scores words by how often they appear in a chunk and how rare they are across all chunks. It needs no download, but it matches words, not meaning. You'll see what that costs in the experiments section.
fit()exists because TF-IDF must learn its vocabulary from your corpus. MiniLM is already trained, sofit()does nothing for it.- Normalisation: the last three lines divide each vector by its length, so a plain dot product later equals cosine similarity. MiniLM's pipeline already ends with a normalisation layer and scikit-learn's TF-IDF normalises by default, so here the step is a safety net. It matters when you swap in an embedding model that does not normalise its output. The
norms[norms == 0] = 1.0line stops a division by zero when a TF-IDF query shares no words with the vocabulary.
Expected output:
== Step 2: embeddings + index ==
backend=minilm index shape=(17, 384) first row norm=1.000
Seventeen chunk vectors, each 384 numbers long and already unit length.
Step 3: The vector index and cosine-similarity search#
def cosine_similarity(a: np.ndarray, b: np.ndarray) -> float:
"""cos(theta) = (a . b) / (||a|| * ||b||)"""
return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
def build_index(chunks: list[dict], embedder: Embedder) -> np.ndarray:
"""Embed every chunk. Row i of the returned matrix is the vector for chunks[i]."""
texts = [c["text"] for c in chunks]
embedder.fit(texts)
return embedder.encode(texts)
def search(query: str, index: np.ndarray, chunks: list[dict], embedder: Embedder, k: int = 3) -> list[dict]:
"""Return the k chunks most similar to the query, best first."""
q = embedder.encode([query])[0]
scores = index @ q # one dot product per chunk (all vectors are unit length)
top = np.argsort(-scores)[:k] # indices of the k highest scores
return [{**chunks[i], "score": float(scores[i])} for i in top]
build_index returns a matrix with one row per chunk; that matrix is our vector store. search embeds the question, computes all scores with a single matrix-vector product (index @ q, the \(\mathbf{s}=E\hat{\mathbf{q}}\) formula above) and sorts. np.argsort(-scores) sorts from highest to lowest, because argsort sorts ascending. Sorting all \(N\) scores is fine for thousands of chunks. For millions you would use an approximate nearest-neighbour index, which is what vector databases provide.
The script first checks the hand-worked cosine example as Step 0, then runs two searches as Step 3:
Expected output:
== Step 0: cosine similarity by hand ==
dot(q,d1)=8 dot(q,d2)=12
cos(q,d1)=0.889 cos(q,d2)=0.667
== Step 3: search ==
Q: I forgot my password. How do I reset it?
0.477 IT-01#0 Passwords and account security
0.273 IT-02#0 Laptops and devices
0.236 IT-02#1 Laptops and devices
Q: Can I use ChatGPT to summarise a customer's complaint?
0.299 SEC-01#0 Data classification and AI tool usage
0.204 HR-03#1 Hybrid and remote work
0.167 FIN-01#0 Expense reimbursement
The Step 0 block matches the hand calculation: 8 and 12 for the dot products, 0.889 and 0.667 for cosine.
The password question shares words with the right chunk, so any method would find it. The ChatGPT question is more interesting: the word "ChatGPT" appears nowhere in the corpus, yet the embedding model ranks the data-classification policy ("never paste Confidential or Restricted data into external AI chatbots") first. That is semantic search at work.
Scores are cosine similarities. Don't read them as probabilities, and don't compare them across different embedding models. Your third decimal place may differ slightly with other library versions or hardware.
Step 4: Augment the prompt#
SYSTEM_PROMPT = (
"You are the HR and IT help assistant for Kestrelwood Systems. "
"Answer ONLY from the context provided. Cite the source id in square brackets, e.g. [HR-01]. "
"If the context does not contain the answer, reply exactly: "
"\"I don't know based on the company documents.\""
)
def build_prompt(question: str, retrieved: list[dict]) -> str:
"""Put the retrieved chunks and the question into one user message."""
context = "\n\n".join(f"[{r['doc_id']}] {r['title']}\n{r['text']}" for r in retrieved)
return f"Context:\n{context}\n\nQuestion: {question}\nAnswer:"
The system prompt does three jobs. It restricts the model to the supplied context, it asks for citations so a user can check the source, and it defines an exact refusal sentence for when the context has no answer, which makes refusals easy to detect and count in testing. Each chunk is labelled with its document id so the model can cite it.
Expected output (top 2 chunks for a leave question):
== Step 4: prompt assembly ==
Context:
[HR-01] Annual leave policy
Every full-time employee at Kestrelwood Systems receives 24 days of paid annual leave per calendar year. Leave accrues at 2 days per month. You can carry forward a maximum of 10 unused days into the next calendar year; any balance above 10 days lapses on 31 December. Apply for leave
[HR-02] Sick leave policy
Employees receive 12 days of paid sick leave per calendar year. Sick leave cannot be carried forward. Inform your manager before 10:00 AM on the day you are unwell. If you are absent for more than 2 consecutive days, upload a medical certificate from a registered doctor to the PeoplePortal
Question: How many vacation days do I get each year?
Answer:
This is precisely the text the LLM will read. When a RAG answer looks wrong, print this prompt first. Most of the time the right chunk is not in it, and the problem is retrieval, not the model.
Step 5: A question the documents cannot answer#
Ask something the corpus cannot answer. The script searches for the share-price question and prints the query itself, then the top hits:
Expected output:
== Step 5: a question the documents cannot answer ==
Q: What is Kestrelwood's share price today?
0.332 HR-01#0 Annual leave policy
0.301 HR-03#0 Hybrid and remote work
0.221 FIN-02#1 Business travel policy
Look at the numbers. The unanswerable share-price question gets a top score of 0.332, higher than the correct chunk for the ChatGPT question (0.299). Retrieval always returns something, and a fixed similarity threshold would either block good answers or let this one through. That's why the refusal instruction in the system prompt matters, and why you test unanswerable questions explicitly.
Step 6 in the script is retrieval evaluation (hit rate@k and recall@k). The full write-up and expected output are in the Evaluate retrieval section below. Next is generation.
Step 7: Generate the answer with an LLM#
Retrieval and prompt assembly are the RAG-specific parts. Generation is a single call to any chat model that accepts a system instruction and a user message. The code keeps that call behind one function, so switching providers never touches retrieval:
def rag_answer(question: str, index: np.ndarray, chunks: list[dict], embedder: Embedder,
k: int = 3, llm_fn=None) -> dict:
"""Retrieve -> augment -> (optionally) generate.
llm_fn is any function that takes (prompt, system_prompt) and returns text,
so you can swap Bedrock for any other chat LLM without touching retrieval.
"""
retrieved = search(question, index, chunks, embedder, k=k)
prompt = build_prompt(question, retrieved)
answer = llm_fn(prompt, SYSTEM_PROMPT) if llm_fn else None
return {"question": question, "retrieved": retrieved, "prompt": prompt, "answer": answer}
Primary example: Amazon Bedrock Converse API (NOT EXECUTED)#
Cost warning: Amazon Bedrock is a paid AWS service. On-demand text models are priced per million input tokens and per million output tokens, and every RAG call sends the retrieved chunks as input tokens. Check Amazon Bedrock pricing for your model and Region before you run this. I did not call Bedrock for this tutorial. The block below was not executed against AWS.
def generate_with_bedrock(prompt: str, system_prompt: str = SYSTEM_PROMPT,
model_id: str = "<YOUR_BEDROCK_MODEL_ID>",
region: str = "us-east-1", client=None) -> str:
"""Send the RAG prompt to a Bedrock model with the Converse API and return the answer text.
COSTS MONEY: each call is billed per input and output token. Requires AWS credentials
with permission for bedrock:InvokeModel. Replace model_id with a model ID (or inference
profile ID) that is enabled in your account and region. Pass `client` to reuse a
boto3 client (or a stubbed one in tests).
"""
if client is None:
import boto3
client = boto3.client("bedrock-runtime", region_name=region)
response = client.converse(
modelId=model_id,
system=[{"text": system_prompt}],
messages=[{"role": "user", "content": [{"text": prompt}]}],
inferenceConfig={"maxTokens": 300, "temperature": 0.0},
)
# The reply is a list of content blocks; keep the text ones (some models add reasoning blocks).
blocks = response["output"]["message"]["content"]
return "".join(block["text"] for block in blocks if "text" in block)
What each part does:
boto3.client("bedrock-runtime", ...): the Converse API is served by thebedrock-runtimeendpoint.modelId: replace the placeholder with a model ID, or an inference profile ID, that is available in your account and Region. Take it from the Bedrock console model catalog or the supported models list; some models are invoked through inference profiles.system: a list of content blocks holding our grounding instructions.messages: the conversation; here oneusermessage whose content is onetextblock containing the RAG prompt.inferenceConfig: the base parameters shared by all models (maxTokens,temperature,topP,stopSequences).temperature=0.0makes the model favour its most likely tokens, which suits factual Q&A.- The answer is in
response["output"]["message"]["content"], a list of content blocks. For a plain text reply that is one block, socontent[0]["text"]works; the code joins every block that has a"text"key because some models also return reasoning blocks. The response also containsstopReasonandusage(input, output and total tokens), which you should log to track cost.
Before the first call you need AWS credentials configured for boto3 (credentials guide) and an IAM identity allowed to perform bedrock:InvokeModel. According to the model access page, access to Bedrock foundation models is enabled by default when the account has the required AWS Marketplace permissions, and Anthropic models additionally need a one-time use-case form. Then run:
python rag_from_scratch.py --generate --model-id <YOUR_MODEL_ID> --region <YOUR_REGION>
How I verified this code without calling AWS. I checked the request fields against the boto3 converse reference and the Bedrock Converse API guide. Then I ran the function against botocore's Stubber, which intercepts the call before it leaves your machine, validates every parameter against the Bedrock Runtime service model that ships with botocore, and returns a fake response. No request reaches AWS and nothing is billed. The check script is verify_bedrock_request.py:
Request parameters valid for Converse: ['inferenceConfig', 'messages', 'modelId', 'system']
Function returned the text blocks of output.message.content -> 'FAKE STUB TEXT - not produced by a model.'
boto3 1.43.102
The "answer" in that output is my placeholder text, not model output. The same validator rejects a wrong parameter name: passing max_tokens instead of maxTokens raised ParamValidationError: Unknown parameter in inferenceConfig: "max_tokens".
Any other chat LLM#
Nothing in the pipeline is Bedrock-specific. Write a function with the signature llm_fn(prompt: str, system_prompt: str) -> str that calls your provider's chat endpoint (OpenAI, Azure OpenAI, Google, a local model server) with the system prompt as the system/instruction message and the RAG prompt as the user message, and pass it as rag_answer(..., llm_fn=your_function). Retrieval, prompt and evaluation stay the same. Check your provider's current SDK documentation for the exact call.
Next step: the same index in FAISS or Chroma#
A NumPy matrix is fine for a few thousand chunks held in memory. A vector library or database adds approximate nearest-neighbour search for millions of vectors, persistence to disk, metadata filtering and incremental updates. Here is the same retrieval in FAISS (a vector search library from Meta) and Chroma (an open-source vector database), reusing our chunks and embeddings (rag_vector_db.py):
"""Next step: the same retrieval with FAISS (a vector index library) and Chroma (a vector database).
Reuses CORPUS, build_chunks and Embedder from rag_from_scratch.py so the results are comparable."""
import faiss
import chromadb
from rag_from_scratch import CORPUS, Embedder, build_chunks, build_index, search
chunks = build_chunks(CORPUS, chunk_size=50, overlap=10)
embedder = Embedder("minilm")
vectors = build_index(chunks, embedder) # (17, 384) float32, unit length
question = "My laptop was stolen at the airport. What should I do?"
q = embedder.encode([question])
print("== NumPy (from scratch) ==")
for r in search(question, vectors, chunks, embedder, k=3):
print(f" {r['score']:.3f} {r['chunk_id']}")
print("== FAISS IndexFlatIP (exact inner product; = cosine because vectors are normalised) ==")
index = faiss.IndexFlatIP(vectors.shape[1])
index.add(vectors)
scores, ids = index.search(q, 3)
for s, i in zip(scores[0], ids[0]):
print(f" {s:.3f} {chunks[i]['chunk_id']}")
print("== Chroma (in-memory client, cosine space, our own embeddings) ==")
client = chromadb.EphemeralClient() # use chromadb.PersistentClient(path=...) to keep data on disk
collection = client.create_collection(name="kestrelwood_policies",
configuration={"hnsw": {"space": "cosine"}})
collection.add(
ids=[c["chunk_id"] for c in chunks],
embeddings=vectors.tolist(),
documents=[c["text"] for c in chunks],
metadatas=[{"doc_id": c["doc_id"], "title": c["title"]} for c in chunks],
)
res = collection.query(query_embeddings=q.tolist(), n_results=3)
for cid, dist in zip(res["ids"][0], res["distances"][0]):
print(f" distance={dist:.3f} similarity={1 - dist:.3f} {cid}")
print("== Chroma with a metadata filter (only FIN-* documents) ==")
res = collection.query(query_embeddings=q.tolist(), n_results=2,
where={"doc_id": {"$in": ["FIN-01", "FIN-02"]}})
print(" ", res["ids"][0])
Expected output:
== NumPy (from scratch) ==
0.504 IT-02#0
0.252 IT-01#0
0.190 FIN-01#0
== FAISS IndexFlatIP (exact inner product; = cosine because vectors are normalised) ==
0.504 IT-02#0
0.252 IT-01#0
0.190 FIN-01#0
== Chroma (in-memory client, cosine space, our own embeddings) ==
distance=0.496 similarity=0.504 IT-02#0
distance=0.748 similarity=0.252 IT-01#0
distance=0.810 similarity=0.190 FIN-01#0
== Chroma with a metadata filter (only FIN-* documents) ==
['FIN-01#0', 'FIN-02#0']
All three agree on the ranking and the scores. Three details:
IndexFlatIPis an exact inner-product index. As the FAISS documentation notes, inner product is cosine similarity only when the vectors are normalised, which ours are.- Chroma's default distance is squared L2, so I set
"space": "cosine"explicitly (Chroma collection configuration). Chroma returns a distance, \(1-\cos\theta\), so smaller is better. - The
wherefilter is something the NumPy version can't do without extra code: restrict search to documents a user is allowed to see, or to one department.
Hands-on lab: extend the assistant and catch a regression#
Allow 30–45 minutes. You need rag_from_scratch.py and the packages above. (A longer guided lab with starter code and TODOs is available for classes: RAG in Python: Student Lab, Cheat Sheet and Quiz.)
Task 1: Baseline. Run python rag_from_scratch.py and confirm your output matches the expected outputs above. Write down hit rate@1 and recall@1 from Step 6 of the output.
Task 2: Add a document and a labelled question. Kestrelwood has published a new holiday policy. Add it with a new evaluation question and re-run the evaluation (lab_task2_add_document.py):
"""Tutorial hands-on lab, Task 2: add a new document and a new labelled question, then re-evaluate."""
from rag_from_scratch import (CORPUS, EVAL_SET, Embedder, build_chunks, build_index,
evaluate_retrieval, search)
new_doc = {
"doc_id": "HR-05",
"title": "Public and floating holidays",
"text": (
"Kestrelwood observes 10 fixed public holidays each calendar year; the list is published "
"in the PeoplePortal every December. In addition, every employee can take 2 floating "
"holidays for festivals of their choice. Apply for a floating holiday in the PeoplePortal "
"like annual leave. Unused floating holidays cannot be carried forward."
),
}
new_question = {"question": "Can I take a day off for a festival that is not on the holiday list?",
"relevant": {"HR-05"}}
corpus = CORPUS + [new_doc]
eval_set = EVAL_SET + [new_question]
chunks = build_chunks(corpus, chunk_size=50, overlap=10)
embedder = Embedder("auto")
index = build_index(chunks, embedder)
print(f"{len(corpus)} documents -> {len(chunks)} chunks")
for r in search(new_question["question"], index, chunks, embedder, k=3):
print(f" {r['score']:.3f} {r['chunk_id']:<9} {r['title']}")
for k in (1, 3):
m = evaluate_retrieval(eval_set, index, chunks, embedder, k=k, verbose=(k == 1))
print(f"k={k}: hit_rate={m['hit_rate']:.2f} recall={m['recall']:.2f}")
Expected output:
11 documents -> 18 chunks
0.540 HR-05#0 Public and floating holidays
0.291 FIN-02#0 Business travel policy
0.281 HR-01#0 Annual leave policy
MISS want=['HR-01'] got=[HR-05] How many vacation days do I get each year?
HIT want=['HR-01'] got=[HR-01] Can I carry unused leave into next year?
HIT want=['HR-02'] got=[HR-02] Do I need a doctor's note if I am ill for three days?
HIT want=['IT-01'] got=[IT-01] I forgot my password. How do I reset it?
HIT want=['IT-02'] got=[IT-02] My laptop was stolen at the airport. What should I do?
HIT want=['FIN-02'] got=[FIN-02] What is the hotel limit when I travel to Mumbai?
HIT want=['SEC-01'] got=[SEC-01] Can I use ChatGPT to summarise a customer's complaint?
HIT want=['HR-04'] got=[HR-04] How much time off do new fathers get?
HIT want=['LND-01'] got=[LND-01] Will the company pay for my AWS certification exam?
HIT want=['FIN-01', 'FIN-02'] got=[FIN-01] How do I book flights for a client visit and claim the costs afterwards?
HIT want=['HR-05'] got=[HR-05] Can I take a day off for a festival that is not on the holiday list?
k=1: hit_rate=0.91 recall=0.86
k=3: hit_rate=1.00 recall=1.00
The new question works: HR-05 is the top result with a clear margin. But hit rate@1 fell from 1.00 to 0.91, because "How many vacation days do I get each year?" now retrieves the holiday policy instead of the annual leave policy. To an embedding model, "vacation" is close to "holidays". Nobody touched that question, and without the evaluation set nobody would have noticed until an employee got the wrong answer.
Task 3: Fix it and prove the fix. Try one change at a time and re-run the evaluation after each:
(a) retrieve two chunks instead of one and check hit rate@2;
(b) prepend each chunk's title to its text before embedding (a common way to add context to short chunks);
(c) add the phrase "annual leave (vacation)" to HR-01, the way real teams add synonyms users actually type.
Record which change restores hit@1 without breaking the new HR-05 question. The point is to decide with numbers, not by trying a few questions by hand.
What I measured (open after you try)
(a) k=2 gives hit rate@2 = 1.00 and recall@2 = 1.00, but it doubles the context sent to the LLM for every question. (b) Prepending titles left hit@1 at 0.91 on this corpus. (c) Adding "(vacation)" after "24 days of paid annual leave" in HR-01 restored hit@1 to 1.00 (recall@1 0.95) with the HR-05 question still correct. On a different corpus the ranking of these fixes could differ, which is exactly why you measure.
Task 4: Feel the difference between TF-IDF and embeddings. Run python rag_from_scratch.py --backend tfidf and compare the evaluation block with the MiniLM one. You should see hit@1 = 0.90 for TF-IDF (the "new fathers" question misses) against 1.00 for MiniLM. The next section explains why.
Common mistakes, with measured effects#
These numbers come from rag_experiments.py, which re-runs the evaluation while changing one setting at a time. Full output:
== Experiment A: embedding backend (chunk_size=50, overlap=10) ==
HIT want=['HR-01'] got=[HR-01] How many vacation days do I get each year?
HIT want=['HR-01'] got=[HR-01] Can I carry unused leave into next year?
HIT want=['HR-02'] got=[HR-02] Do I need a doctor's note if I am ill for three days?
HIT want=['IT-01'] got=[IT-01] I forgot my password. How do I reset it?
HIT want=['IT-02'] got=[IT-02] My laptop was stolen at the airport. What should I do?
HIT want=['FIN-02'] got=[FIN-02] What is the hotel limit when I travel to Mumbai?
HIT want=['SEC-01'] got=[SEC-01] Can I use ChatGPT to summarise a customer's complaint?
MISS want=['HR-04'] got=[IT-02] How much time off do new fathers get?
HIT want=['LND-01'] got=[LND-01] Will the company pay for my AWS certification exam?
HIT want=['FIN-01', 'FIN-02'] got=[FIN-02] How do I book flights for a client visit and claim the costs afterwards?
tfidf hit@1=0.90 recall@1=0.85 | hit@3=1.00 recall@3=1.00
minilm hit@1=1.00 recall@1=0.95 | hit@3=1.00 recall@3=1.00
== Experiment B: chunk size and overlap (MiniLM) ==
size=12 overlap=0 chunks=49 hit@1=0.90 recall@1=0.90 recall@3=0.95
size=12 overlap=4 chunks=69 hit@1=0.90 recall@1=0.90 recall@3=0.95
size=50 overlap=10 chunks=17 hit@1=1.00 recall@1=0.95 recall@3=1.00
size=500 overlap=0 chunks=10 hit@1=1.00 recall@1=0.95 recall@3=1.00
== Experiment C: what a 12-word chunk with no overlap looks like ==
HR-02#0: Employees receive 12 days of paid sick leave per calendar year. Sick
HR-02#1: leave cannot be carried forward. Inform your manager before 10:00 AM on
HR-02#2: the day you are unwell. If you are absent for more than
HR-02#3: 2 consecutive days, upload a medical certificate from a registered doctor to
HR-02#4: the PeoplePortal within 3 days of returning to work.
== Experiment D: stuffing more chunks into the prompt (MiniLM, size=50, overlap=10) ==
k=1 precision@k=1.00 prompt length for Q1=66 words
k=3 precision@k=0.53 prompt length for Q1=156 words
k=5 precision@k=0.32 prompt length for Q1=228 words
k=10 precision@k=0.18 prompt length for Q1=467 words
== Experiment E: does the top chunk contain the answer? (MiniLM) ==
size=12 overlap=0: 0.537 HR-04#0: Birth mothers receive 26 weeks of paid maternity leave. Fathers, partners and
size=12 overlap=4: 0.668 HR-04#1: leave. Fathers, partners and adoptive parents receive 4 weeks of paid parental
1. Chunks that are too small. At 12 words per chunk (Experiment B), hit@1 drops to 0.90. Experiment C shows why: the sick-leave rule is shredded into fragments like "the day you are unwell. If you are absent for more than", and a fragment carries too little meaning to match a question well. The query that failed here was the two-document question about booking flights and claiming costs.
2. Chunks that are too large. In this corpus, whole documents (size=500) score the same as 50-word chunks, because every document is 48–72 words long. Don't generalise from that. With real multi-page documents, one vector per document blurs many topics together, anything past the model's input limit (256 word pieces for MiniLM) is not embedded at all, and each retrieved "chunk" fills the prompt with mostly irrelevant text. Start around a paragraph and tune with your evaluation set.
3. No overlap. Overlap didn't change the document-level scores in Experiment B, but Experiment E shows what those scores hide. For "How much time off do new fathers get?", the 12-word no-overlap run retrieves the correct document, yet its top chunk ends at "Fathers, partners and", so the answer ("4 weeks") is not in it. With 4 words of overlap, the top chunk contains "Fathers, partners and adoptive parents receive 4 weeks of paid parental", and it scores higher (0.668 against 0.537). Document-level hit rate counted both runs as a hit. Only one of them would let the LLM answer correctly.
4. Wrong similarity metric or unnormalised vectors. The hand example showed raw dot product preferring a vector only because it was longer. If your embedding model doesn't normalise its output, either normalise before indexing or use cosine explicitly. In a vector database, check the default: Chroma defaults to squared L2. For normalised vectors L2 gives the same ranking as cosine (the \(2-2\cos\theta\) identity), but for unnormalised vectors it doesn't. Also use the same embedding model for documents and queries. Vectors from two different models are not comparable.
5. Stuffing too many chunks into the prompt. Experiment D: at k=1, every retrieved chunk is from a relevant document (precision 1.00); at k=10 only 18% are, and the prompt is 7 times longer (467 against 66 words). Irrelevant context costs input tokens on every call and gives the model more material to get confused by. Retrieve a little more than you need, then keep only the best few, or add a reranker.
6. Relying on keyword matching for paraphrased questions. In Experiment A, TF-IDF misses "How much time off do new fathers get?": after stop-word removal the query terms are "time", "new" and "fathers" (plus two-word pairs), and the laptop policy wins on "new" ("Every new joiner...") with a score of 0.087 against 0.070 for the parental-leave policy's "fathers". Embeddings match meaning. Keyword search still has a place for exact codes, product names and error numbers, which is why many production systems combine both (hybrid search).
7. Shipping without an evaluation set. Every finding above, including the regression in the lab, came from ten labelled questions. Without them you would be judging your system by trying a few questions by hand and reading the answers.
Evaluate retrieval: hit rate@k and recall@k#
Evaluate retrieval separately from generation. If the right chunk isn't retrieved, no model can answer correctly, and retrieval metrics are cheap and deterministic, with no LLM needed.
For each question \(q\) in the question set \(Q\), let \(R_q\) be the set of documents that contain the answer (your labels) and \(T_q(k)\) the set of documents behind the top \(k\) retrieved chunks.
- \(|Q|\) is the number of questions.
- \(\mathbf{1}[\cdot]\) is the indicator function: 1 if the condition is true, 0 otherwise.
- \(R_q\cap T_q(k)\neq\varnothing\) means "at least one relevant document was retrieved".
- \(|R_q\cap T_q(k)|\) is how many of the relevant documents were retrieved; \(|R_q|\) is how many there are.
Hit rate asks "did we find something useful?"; recall asks "did we find everything needed?". They differ only for questions with more than one relevant document, like our flights-and-expenses question.
def evaluate_retrieval(eval_set: list[dict], index: np.ndarray, chunks: list[dict],
embedder: Embedder, k: int = 3, verbose: bool = False) -> dict:
"""Hit rate@k: share of questions with at least one relevant doc in the top k.
Recall@k: average share of each question's relevant docs found in the top k."""
hits, recalls = [], []
for item in eval_set:
results = search(item["question"], index, chunks, embedder, k=k)
found = {r["doc_id"] for r in results} & item["relevant"]
hits.append(1.0 if found else 0.0)
recalls.append(len(found) / len(item["relevant"]))
if verbose:
got = ", ".join(r["doc_id"] for r in results)
mark = "HIT " if found else "MISS"
print(f" {mark} want={sorted(item['relevant'])} got=[{got}] {item['question']}")
return {"k": k, "hit_rate": float(np.mean(hits)), "recall": float(np.mean(recalls))}
Expected output (MiniLM, chunk_size=50, overlap=10):
== Step 6: retrieval evaluation ==
HIT want=['HR-01'] got=[HR-01] How many vacation days do I get each year?
HIT want=['HR-01'] got=[HR-01] Can I carry unused leave into next year?
HIT want=['HR-02'] got=[HR-02] Do I need a doctor's note if I am ill for three days?
HIT want=['IT-01'] got=[IT-01] I forgot my password. How do I reset it?
HIT want=['IT-02'] got=[IT-02] My laptop was stolen at the airport. What should I do?
HIT want=['FIN-02'] got=[FIN-02] What is the hotel limit when I travel to Mumbai?
HIT want=['SEC-01'] got=[SEC-01] Can I use ChatGPT to summarise a customer's complaint?
HIT want=['HR-04'] got=[HR-04] How much time off do new fathers get?
HIT want=['LND-01'] got=[LND-01] Will the company pay for my AWS certification exam?
HIT want=['FIN-01', 'FIN-02'] got=[FIN-01] How do I book flights for a client visit and claim the costs afterwards?
k=1: hit_rate=1.00 recall=0.95
k=3: hit_rate=1.00 recall=1.00
At k=1, every question finds a relevant document (hit rate 1.00), but the two-document question can only get one of its two documents from a single chunk, so recall is 0.95. At k=3 both reach 1.00.
Keep perspective: ten questions over ten short documents is a teaching set, and a perfect score on it proves little. For a real system, collect 50–200 questions from actual users, label the documents that answer them, include questions that should be refused, and re-run the evaluation on every change to chunking, embedding model or corpus. After retrieval, evaluate the generated answers too (faithfulness to the context, answer correctness). The RAGAS paper by Es et al. (2023) is a good starting point for that.
Real-world applications#
- HR and IT helpdesks. Our example: policy questions answered with a link to the exact clause, which cuts down repeat tickets.
- Customer support. Answers over product manuals, release notes and resolved tickets, with citations the agent can check before replying.
- Engineering runbooks. On-call engineers ask "how do we fail over the reporting database?" and get the current runbook section rather than a wiki search results page.
- Compliance and contracts. Finding the clauses that match a question across many agreements, filtered by client or date through metadata.
- Education. A course tutor that answers only from the course material and cites the lesson, which is the idea behind my "Build a RAG tutor" project.
- Agents. A RAG lookup is often the first tool an AI agent gets. See where RAG sits on the autonomy ladder in Rise of the AI Agents.
Training a team to build and evaluate RAG on your own documents and stack? See RAG & LLM Applications training or message me on WhatsApp.
Interview questions and answers#
1. What problem does RAG solve, and what doesn't it solve?
It gives an LLM access to private or recent information at question time by retrieving relevant passages and adding them to the prompt, so answers are grounded and citable without retraining. It doesn't guarantee correctness: if retrieval misses, or the model ignores or misreads the context, the answer can still be wrong. It also doesn't change the model's reasoning ability or style.
2. RAG or fine-tuning: how do you choose?
Use RAG when the knowledge changes often, must be cited, or must respect per-user access. Use fine-tuning to change behaviour, format or tone, or to teach a narrow task. They combine well: a fine-tuned model can still use retrieved context. Updating a document in a RAG system takes a re-index; updating facts through fine-tuning takes another training run.
3. Why do we chunk documents, and how do you choose chunk size?
One embedding per long document averages many topics into one vague vector, and embedding models truncate long input. Chunks keep each vector specific. Size is a trade-off: too small loses context (our 12-word chunks dropped hit@1 to 0.90), too large dilutes the vector and fills the prompt. Start around a paragraph with some overlap, respect sentence and heading boundaries, and tune with a labelled evaluation set.
4. Explain cosine similarity and why embeddings are often normalised.
Cosine similarity is the dot product divided by the product of the vector lengths, so it measures the angle and ignores length. If you normalise every vector to length 1 once at indexing time, the dot product equals the cosine, so search becomes one matrix-vector multiply, and L2-distance ranking gives the same order as cosine.
5. How do you evaluate a RAG system?
In two layers. Retrieval: hit rate@k, recall@k (and MRR or nDCG if rank order matters) on hand-labelled questions. Generation: faithfulness to the retrieved context, answer correctness and refusal behaviour on unanswerable questions, checked by humans or an LLM judge with spot checks. Re-run both on every change, as a regression test.
6. A user reports a wrong answer. How do you debug it?
Log and inspect the retrieved chunks and the exact prompt. If the correct chunk wasn't retrieved, it's a retrieval problem: chunking, embedding model, query wording, top-k, missing metadata filter, or the document isn't indexed. If it was retrieved but the answer is wrong, it's a generation problem: prompt instructions, too much irrelevant context, or the model. Add the question to the evaluation set so it stays fixed.
7. What does a vector database add over a NumPy array?
Approximate nearest-neighbour indexes (such as HNSW) that search millions of vectors in milliseconds, persistence, incremental inserts and deletes, metadata filtering (for example per-user access control), and operational features such as replication and backups. For a few thousand chunks in one process, exact search in NumPy is simpler and fully accurate.
8. What is hybrid search and when would you use it?
Combining keyword scoring (BM25 or TF-IDF) with vector similarity and merging the ranked results. Embeddings handle paraphrases, while keywords handle exact identifiers such as error codes, SKUs and names that embeddings can blur. Our experiment shows each approach failing differently: TF-IDF missed "new fathers" on a word-overlap technicality, and embeddings confused "vacation" with "holidays".
Test yourself: 10 MCQs#
Q1. What is the main reason an LLM on its own gives wrong answers about your company's internal policies?A) Its context window is too smallB) It was never trained on those documents and still produces a fluent answerC) Its temperature is too highD) It cannot read English policies
Q2. In RAG, which step turns text into vectors?A) ChunkingB) EmbeddingC) AugmentationD) Generation
Q3. For \(\mathbf{a}=(3,4)\) and \(\mathbf{b}=(4,3)\), what is \(\cos\theta\)?A) 0.50B) 0.75C) 0.96D) 1.00
Q4. All vectors in an index have length 1. Which statement is true?A) Dot product equals cosine similarityB) Cosine similarity is always 1C) L2 distance becomes meaninglessD) You must use Manhattan distance
Q5. With chunk_size=50 and overlap=10, at which word does the second chunk start?A) 10B) 40C) 50D) 60
Q6. Why add overlap between chunks?A) To make the index smallerB) To reduce the chance that a fact is split across a chunk boundaryC) To speed up embeddingD) To remove duplicate documents
Q7. A question has two relevant documents and the top-3 results contain one of them. For this question, what are hit@3 and recall@3?A) 1 and 1B) 1 and 0.5C) 0.5 and 0.5D) 0 and 0.5
Q8. In our experiment, raising k from 1 to 10 reduced precision@k from 1.00 to 0.18. What is the main practical risk of a large k?A) The LLM refuses to answerB) Higher cost and more irrelevant context competing with the right chunkC) The embeddings changeD) Recall goes down
Q9. In the Bedrock Converse API response, where is the generated text?A) response["body"]["completion"]B) response["output"]["message"]["content"][0]["text"]C) response["choices"][0]["text"]D) response["generation"]
Q10. Which retrieval method is most likely to match "Can I use ChatGPT with customer data?" to a policy that says "never paste Restricted data into external AI chatbots"?A) Exact string matchB) TF-IDF with stop words removedC) Dense embeddings with cosine similarityD) Sorting documents by length
Answers and explanations#
Show answers and explanations
- B. The documents weren't in its training data, and next-token prediction produces a plausible answer anyway. Context size and temperature affect other behaviour but aren't the root cause.
- B. The embedding model maps each chunk (and each question) to a vector. Chunking splits text, augmentation builds the prompt, generation produces the answer.
- C. \(\mathbf{a}\cdot\mathbf{b}=12+12=24\), \(\lVert\mathbf{a}\rVert=\lVert\mathbf{b}\rVert=5\), so \(\cos\theta=24/25=0.96\).
- A. The cosine denominator becomes \(1\times1\). L2 distance still works, and ranks in the same order as cosine because \(\lVert\hat{a}-\hat{b}\rVert^2=2-2\cos\theta\).
- B. The window moves by
chunk_size - overlap= 40 words, so chunks start at 0, 40, 80, and so on. - B. Overlap repeats boundary words so a fact cut at a boundary appears whole in at least one chunk. Experiment E showed the top chunk gaining the "4 weeks" fact.
- B. Hit counts whether any relevant document was found (yes = 1); recall is the fraction found (1 of 2 = 0.5).
- B. More chunks mean more input tokens on every call and more distracting text. Recall can only stay the same or rise as k grows.
- B. Converse returns
output.message.content, a list of content blocks; the first text block holds the answer. The other options are not fields of the Converse response. - C. "ChatGPT" never appears in the policy, so exact match and TF-IDF have little to work with. Our MiniLM search ranked SEC-01 first for this question.
Practice exercises#
- Sentence-aware chunking. Rewrite
chunk_textto split on sentence boundaries (for example with a regular expression on.), packing whole sentences until a word limit. Compare hit@1 and Experiment E's top chunk with the word-based version. - MRR. Add mean reciprocal rank to
evaluate_retrieval: for each question, 1/rank of the first relevant document (0 if none), averaged. Report it for k=3. - Deduplicate overlapping chunks. When two retrieved chunks come from the same document and overlap, merge them before building the prompt. Measure the prompt length before and after.
- Refusal test. Add five unanswerable questions to a separate list. Log the top score for each and plot them next to the top scores of the answerable questions. Can any threshold separate the two groups on this corpus?
- Hybrid search. Compute both TF-IDF and MiniLM scores, combine them as \(0.5\,s_{\text{tfidf}} + 0.5\,s_{\text{minilm}}\), and see whether the "new fathers" TF-IDF miss and the lab's "vacation" regression both disappear.
- Persist the index. Change
rag_vector_db.pyto usechromadb.PersistentClient(path="kestrelwood_db"), run it twice, and confirm the second run can query without re-adding the chunks. - Generation (optional, billed). If you have Bedrock access, run
--generatefor the ten evaluation questions plus the share-price question. Count how many answers cite the right document id and whether the share-price question returns the exact refusal sentence.
Summary#
- An LLM on its own can't know your private or recent information and will still answer. RAG retrieves relevant passages at question time and puts them in the prompt.
- The pipeline is ingest → chunk → embed → store → retrieve → augment → generate. The first four run when documents change; the last three run per question.
- Cosine similarity measures the angle between embedding vectors. Normalise once and it becomes a dot product, so search is one matrix-vector multiply.
- Working retrieval (chunking, embeddings, index, search) takes under 100 lines of Python and NumPy. FAISS and Chroma give identical results here and add scale, persistence and metadata filtering.
- Generation is one pluggable call. The Bedrock Converse example is ready to run but costs money, so it is behind a flag.
- Measure retrieval with hit rate@k and recall@k on labelled questions. In this tutorial those numbers exposed too-small chunks, a hidden overlap problem, TF-IDF's keyword blind spot, over-stuffed prompts and a regression caused by adding one document.
Further learning#
- Patrick Lewis et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020. https://arxiv.org/abs/2005.11401
- Vladimir Karpukhin et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering (the retriever used in the RAG paper). https://arxiv.org/abs/2004.04906
- Nils Reimers and Iryna Gurevych (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. https://arxiv.org/abs/1908.10084
- Model card for
all-MiniLM-L6-v2(384 dimensions, 256 word-piece input limit). https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 - Sentence Transformers documentation, Semantic Search. https://sbert.net/examples/sentence_transformer/applications/semantic-search/README.html
- Amazon Bedrock User Guide, Inference using the Converse API. https://docs.aws.amazon.com/bedrock/latest/userguide/conversation-inference.html
- Boto3 reference,
BedrockRuntime.Client.converse. https://boto3.amazonaws.com/v1/documentation/api/latest/reference/services/bedrock-runtime/client/converse.html - Amazon Bedrock pricing. https://aws.amazon.com/bedrock/pricing/
- FAISS wiki, MetricType and distances. https://github.com/facebookresearch/faiss/wiki/MetricType-and-distances
- Chroma documentation, Configure collections. https://docs.trychroma.com/docs/collections/configure
- Christopher Manning, Prabhakar Raghavan and Hinrich Schütze. Introduction to Information Retrieval, Chapter 8: Evaluation in information retrieval. https://nlp.stanford.edu/IR-book/html/htmledition/evaluation-in-information-retrieval-1.html
- Shahul Es et al. (2023). RAGAS: Automated Evaluation of Retrieval Augmented Generation. https://arxiv.org/abs/2309.15217
Next on this site: the foundations behind this tutorial in AI & Machine Learning Foundations and what happens when retrieval becomes one tool among many in Rise of the AI Agents. Want to build this with me live, over real documents? Join the RAG & LLM Applications batch or message me on WhatsApp.
Continue
Practise it: the companion student lab, cheat sheet and 15-question quiz uses the same corpus and function names.
Learn it live: RAG & LLM Applications training · WhatsApp +91 70492 35525