Building a Basic RAG Pipeline
Learning Outcomes
- Ingest documents from PDFs and plain text into a normalised Python structure
- Chunk and embed documents using
langchain-text-splittersandsentence-transformers - Store and query embeddings in a local Chroma vector database
- Assemble retrieved context into a grounded prompt and generate a cited answer with the Anthropic SDK
- Verify the pipeline end-to-end against a set of known-answer test queries
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Pipeline overview and what we're building |
| Step 1–2 | 10 min | Ingestion and text cleaning |
| Step 3–4 | 12 min | Chunking and embedding |
| Step 5 | 10 min | Vector storage with Chroma |
| Step 6 | 10 min | Retrieval and context assembly |
| Step 7–8 | 10 min | Grounded generation and smoke-testing |
| Wrap-up | 5 min | Key takeaways and what's next |
Before You Begin
Pre-work:
- Read Lesson 2: Embeddings Explained — you need to understand cosine similarity before embedding documents
- Read Lesson 3: Vector Databases — Chroma concepts are used directly here
- Read Lesson 4: Chunking Strategies — we apply recursive splitting and discuss why in context
Shopping List:
- Python 3.10 or later
- An Anthropic API key set as the
ANTHROPIC_API_KEYenvironment variable - The following packages installed (one
pip installshown in Step 1):pypdf,langchain-text-splitters,sentence-transformers,chromadb,anthropic - A small corpus of documents to index — two or three PDFs or
.txtfiles work fine for this lesson
Start from a clean project folder. The layout below keeps pipeline stages clearly separated so you can swap components later without touching unrelated code:
rag-pipeline/
├── data/ # raw source documents (PDFs, .txt, .md)
├── ingest.py # load + clean + chunk + embed + store
├── retrieve.py # embed query, similarity search, assemble context
├── generate.py # call the LLM with grounded prompt
├── pipeline.py # end-to-end runner
└── smoke_test.py # verify the pipeline against known queries
Install all runtime dependencies in one command:
pip install pypdf langchain-text-splitters sentence-transformers chromadb anthropic
Verified package names as of 2026-05-31:
pypdf— pure-Python PDF readerlangchain-text-splitters— standalone text splitting utilities from LangChainsentence-transformers— embedding and reranker models from Hugging Facechromadb— open-source vector database with a Python-native APIanthropic— official Anthropic Python SDK
Pin these to a requirements.txt for reproducibility:
pip freeze | grep -E "pypdf|langchain|sentence|chromadb|anthropic" > requirements.txt
python -m venv .venv && source .venv/bin/activate. The sentence-transformers wheel downloads model weights on first use, so allow a few minutes and a few hundred MB of disk space the first time.The first stage converts raw files into normalised Python dicts. Each dict carries the text content plus metadata fields you will later store alongside embeddings in Chroma.
# ingest.py -- Part 1: load raw documents
import re
from pathlib import Path
from pypdf import PdfReader
def load_pdf(path: Path) -> dict:
"""Extract text from every page of a PDF and return a doc dict."""
reader = PdfReader(str(path))
pages = []
for page_num, page in enumerate(reader.pages):
raw = page.extract_text() or ""
pages.append({"page": page_num + 1, "text": raw})
full_text = "\n\n".join(p["text"] for p in pages)
return {
"source": str(path),
"filename": path.name,
"type": "pdf",
"text": full_text,
"pages": len(reader.pages),
}
def load_text(path: Path) -> dict:
"""Load a plain-text or Markdown file."""
text = path.read_text(encoding="utf-8")
return {
"source": str(path),
"filename": path.name,
"type": "text",
"text": text,
"pages": 1,
}
def load_corpus(data_dir: str = "data") -> list[dict]:
"""Load all supported files from a directory."""
corpus = []
for path in sorted(Path(data_dir).iterdir()):
if path.suffix.lower() == ".pdf":
corpus.append(load_pdf(path))
elif path.suffix.lower() in {".txt", ".md"}:
corpus.append(load_text(path))
return corpus
After loading, run a lightweight cleaning pass to remove noise that degrades embedding quality:
# ingest.py -- Part 2: text cleaning
def clean_text(text: str) -> str:
"""Remove common PDF artefacts and normalise whitespace."""
# Collapse runs of whitespace and blank lines
text = re.sub(r"[ \t]+", " ", text)
text = re.sub(r"\n{3,}", "\n\n", text)
# Remove page-number lines (lines that are only a digit)
text = re.sub(r"(?m)^\s*\d+\s*$", "", text)
# Remove repeated header/footer patterns (heuristic: same line 3+ times)
lines = text.split("\n")
from collections import Counter
line_counts = Counter(ln.strip() for ln in lines if ln.strip())
boilerplate = {ln for ln, cnt in line_counts.items() if cnt >= 3 and len(ln) < 120}
cleaned = [ln for ln in lines if ln.strip() not in boilerplate]
return "\n".join(cleaned).strip()
The cleaning heuristics above handle the most common issues in digitised PDFs: duplicate headers, page numbers, and excessive whitespace. For scanned PDFs you will need an OCR layer (pdfplumber or Tesseract) — that is outside scope here.
page.extract_text() output before indexing — if you get back empty strings your document needs OCR first.| Source type | Load function | Notes |
|---|---|---|
| PDF (digital) | load_pdf() |
Uses pypdf PdfReader |
| Plain text / Markdown | load_text() |
UTF-8 read |
| Web page | requests.get() + BeautifulSoup |
Not shown here; same dict shape |
source, filename, and page (or section) on every chunk. You cannot cite a source at generation time if you haven't tracked it through the pipeline. Good metadata is what separates a demo from a production system.Chunking converts a long document text into retrieval units. The Lesson 4 deep-dive covers all major strategies; here we use RecursiveCharacterTextSplitter — the best default for prose documents.
It tries to split on \n\n first (paragraph boundary), then \n (line), then space, so chunk boundaries naturally align with sentence and paragraph edges rather than cutting mid-sentence.
# ingest.py -- Part 3: chunking
from langchain_text_splitters import RecursiveCharacterTextSplitter
CHUNK_SIZE = 512 # characters, not tokens
CHUNK_OVERLAP = 64 # carry the last 64 chars into the next chunk
def chunk_document(doc: dict) -> list[dict]:
"""Split a document into overlapping text chunks, preserving metadata."""
splitter = RecursiveCharacterTextSplitter(
chunk_size=CHUNK_SIZE,
chunk_overlap=CHUNK_OVERLAP,
separators=["\n\n", "\n", ". ", " ", ""],
)
texts = splitter.split_text(clean_text(doc["text"]))
chunks = []
for i, text in enumerate(texts):
chunks.append({
"id": f"{doc['filename']}-chunk-{i:04d}",
"text": text,
"source": doc["source"],
"filename": doc["filename"],
"chunk_index": i,
"total_chunks": len(texts),
})
return chunks
def chunk_corpus(docs: list[dict]) -> list[dict]:
all_chunks = []
for doc in docs:
all_chunks.extend(chunk_document(doc))
print(f"Chunked {len(docs)} documents into {len(all_chunks)} chunks.")
return all_chunks
Chunk size trade-offs are worth internalising before you tune:
| Chunk size | Retrieval precision | Context richness | Risk |
|---|---|---|---|
| Small (128–256 chars) | High — narrow match | Low — loses surrounding context | Answer lacks detail |
| Medium (512–800 chars) | Balanced | Balanced | Best starting point |
| Large (1 200+ chars) | Low — too much noise | High — rich context | Irrelevant content dilutes the prompt |
The 64-character overlap carries sentence fragments across boundaries, preventing the retriever from missing a fact that straddles a chunk edge.
An embedding is a dense vector that encodes the meaning of a piece of text. Mathematically similar texts produce vectors that are close together (high cosine similarity). This is how the retriever finds relevant chunks without keyword matching.
We use sentence-transformers with the all-MiniLM-L6-v2 model — 384-dimensional vectors, fast CPU inference, strong English-language retrieval, and no API call required.
# ingest.py -- Part 4: embedding
from sentence_transformers import SentenceTransformer
EMBEDDING_MODEL = "sentence-transformers/all-MiniLM-L6-v2"
# Load once; model weights are cached after the first download
_model: SentenceTransformer | None = None
def get_model() -> SentenceTransformer:
global _model
if _model is None:
_model = SentenceTransformer(EMBEDDING_MODEL)
return _model
def embed_chunks(chunks: list[dict]) -> list[dict]:
"""Add an 'embedding' key (list of floats) to each chunk dict."""
model = get_model()
texts = [c["text"] for c in chunks]
# encode() returns a numpy array of shape (N, 384)
vectors = model.encode(texts, batch_size=64, show_progress_bar=True)
for chunk, vector in zip(chunks, vectors):
chunk["embedding"] = vector.tolist() # Chroma expects a plain list
return chunks
A few things to note about this implementation:
batch_size=64processes 64 chunks per forward pass, which is efficient on CPU. Raise to 128 or 256 if you have a GPU.encode()returnsnumpy.ndarray; Chroma's Python client expectslist[float], so we call.tolist().- The model is loaded once at import time and reused — never reinstantiate per chunk or per query.
Embedding the query at retrieval time must use the same model. Using different models for documents and queries is one of the most common RAG bugs and produces near-random retrieval quality.
all-MiniLM-L6-v2 is a good default. For multilingual content, paraphrase-multilingual-MiniLM-L12-v2 covers 50+ languages. For the highest-quality English retrieval, check the MTEB leaderboard for current top models — it is updated continuously.Chroma is an open-source vector database with a Python-native API and no infrastructure dependencies for local use. PersistentClient writes to disk so your index survives restarts without re-embedding.
# ingest.py -- Part 5: store in Chroma
import chromadb
CHROMA_PATH = "./chroma_db"
COLLECTION_NAME = "rag_docs"
def get_collection(reset: bool = False) -> chromadb.Collection:
client = chromadb.PersistentClient(path=CHROMA_PATH)
if reset:
try:
client.delete_collection(COLLECTION_NAME)
except Exception:
pass
collection = client.get_or_create_collection(
name=COLLECTION_NAME,
metadata={"hnsw:space": "cosine"}, # use cosine distance
)
return collection
def store_chunks(chunks: list[dict], reset: bool = False) -> None:
"""Upsert all chunks (with embeddings) into the Chroma collection."""
collection = get_collection(reset=reset)
ids = [c["id"] for c in chunks]
embeddings = [c["embedding"] for c in chunks]
documents = [c["text"] for c in chunks]
metadatas = [
{
"source": c["source"],
"filename": c["filename"],
"chunk_index": c["chunk_index"],
}
for c in chunks
]
# Chroma's add() is idempotent when IDs already exist (upsert behaviour)
collection.add(
ids=ids,
embeddings=embeddings,
documents=documents,
metadatas=metadatas,
)
print(f"Stored {len(chunks)} chunks. Collection count: {collection.count()}")
The "hnsw:space": "cosine" metadata key tells Chroma to build the HNSW index using cosine distance rather than L2 (Euclidean). Cosine distance is the right choice for sentence embeddings because it measures directional similarity, which is what the embedding model was trained to optimise.
| Chroma client | Storage | When to use |
|---|---|---|
EphemeralClient() |
RAM only | Unit tests, CI |
PersistentClient(path=...) |
Local disk | Development, single-machine production |
HttpClient(host=..., port=...) |
Remote server | Multi-process, containerised |
Now wire up the full ingestion pipeline:
# ingest.py -- Part 6: main ingestion runner
def run_ingestion(data_dir: str = "data", reset: bool = False) -> None:
docs = load_corpus(data_dir)
chunks = chunk_corpus(docs)
chunks = embed_chunks(chunks)
store_chunks(chunks, reset=reset)
print("Ingestion complete.")
if __name__ == "__main__":
run_ingestion(reset=True)
Run it:
python ingest.py
You should see progress output from encode() and a final count. If you rerun without reset=True, add() is idempotent — existing IDs are updated, new IDs are inserted.
get_or_create_collection() without reset=True and call collection.add() with only the new chunks. Track which files are already indexed (e.g. by storing a hash in a sidecar SQLite DB) to avoid re-embedding unchanged files.Retrieval is the core of RAG: embed the user's question, find the nearest document chunks, and assemble them into a context block the LLM can reason over.
# retrieve.py
import chromadb
from sentence_transformers import SentenceTransformer
from ingest import CHROMA_PATH, COLLECTION_NAME, EMBEDDING_MODEL
TOP_K = 5 # number of chunks to retrieve
def retrieve(query: str, top_k: int = TOP_K) -> list[dict]:
"""Embed query, search Chroma, return ranked results."""
model = SentenceTransformer(EMBEDDING_MODEL)
query_vector = model.encode(query).tolist()
client = chromadb.PersistentClient(path=CHROMA_PATH)
collection = client.get_or_create_collection(
name=COLLECTION_NAME,
metadata={"hnsw:space": "cosine"},
)
results = collection.query(
query_embeddings=[query_vector],
n_results=top_k,
include=["documents", "metadatas", "distances"],
)
hits = []
for doc, meta, dist in zip(
results["documents"][0],
results["metadatas"][0],
results["distances"][0],
):
hits.append({
"text": doc,
"source": meta.get("source", "unknown"),
"filename": meta.get("filename", "unknown"),
"chunk_index": meta.get("chunk_index", -1),
"score": 1.0 - dist, # convert distance to similarity (0..1)
})
return hits
def assemble_context(hits: list[dict]) -> str:
"""Format retrieved chunks into a numbered context block."""
parts = []
for i, hit in enumerate(hits, start=1):
parts.append(
f"[Source {i}: {hit['filename']}, chunk {hit['chunk_index']}]\n"
f"{hit['text']}"
)
return "\n\n---\n\n".join(parts)
The score field (1 - cosine distance) gives you a relevance signal in [0, 1]. You can use it to filter low-quality hits before assembly:
MIN_SCORE = 0.30
def retrieve_filtered(query: str, top_k: int = TOP_K) -> list[dict]:
hits = retrieve(query, top_k=top_k * 2) # over-fetch, then filter
return [h for h in hits if h["score"] >= MIN_SCORE][:top_k]
Over-fetching then filtering is a practical pattern for skipping the noise chunks that fall below the relevance threshold without tuning n_results per query.
retrieve() function loads the same EMBEDDING_MODEL constant defined in ingest.py. Never hardcode the model name in two places — always import from a single source-of-truth constant. A mismatch here is invisible at startup but produces retrieval scores near 0.The final stage passes the assembled context to Claude alongside the user's question. The system prompt instructs Claude to ground its answer strictly in the retrieved context and to cite the source labels introduced in assemble_context().
# generate.py
import os
import anthropic
from retrieve import retrieve_filtered, assemble_context
SYSTEM_PROMPT = """You are a precise research assistant. Answer the user's question
using ONLY the information in the provided context sources. Rules:
1. Base every factual claim on a specific source. Cite it inline as [Source N].
2. If the context does not contain enough information to answer, say so explicitly.
Do NOT fabricate facts or draw on general knowledge.
3. Be concise. Prefer one clear paragraph over a verbose list unless a list is
genuinely clearer.
4. If two sources conflict, acknowledge the conflict and cite both."""
GENERATION_MODEL = "claude-sonnet-4-6"
def generate(query: str, top_k: int = 5) -> dict:
"""Full RAG cycle: retrieve -> assemble -> generate."""
hits = retrieve_filtered(query, top_k=top_k)
if not hits:
return {
"answer": "No relevant documents found for this query.",
"sources": [],
"hits": [],
}
context = assemble_context(hits)
user_message = f"""Context sources:
{context}
---
Question: {query}
Answer (cite sources inline):"""
client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
response = client.messages.create(
model=GENERATION_MODEL,
max_tokens=1024,
system=SYSTEM_PROMPT,
messages=[{"role": "user", "content": user_message}],
)
answer = response.content[0].text
sources = [
{"label": f"Source {i+1}", "filename": h["filename"], "score": h["score"]}
for i, h in enumerate(hits)
]
return {"answer": answer, "sources": sources, "hits": hits}
Note the structure of the prompt:
- The system prompt sets hard constraints (cite sources, no fabrication, acknowledge gaps).
- The user message embeds the numbered context block then asks the question below it. Separating context from question with a horizontal rule gives the model a clear boundary.
- The cited
[Source N]labels map directly to the filenames insources, so you can present them in a UI.
Wire everything into pipeline.py:
# pipeline.py
from generate import generate
import json
if __name__ == "__main__":
query = input("Question: ").strip()
result = generate(query)
print("\n=== Answer ===")
print(result["answer"])
print("\n=== Sources used ===")
for s in result["sources"]:
print(f" {s['label']}: {s['filename']} (relevance {s['score']:.2f})")
Run it:
python pipeline.py
stream=True to client.messages.create() and iterate over the stream to print tokens as they arrive. The Anthropic SDK supports streaming natively — see the SDK docs for the stream() context manager pattern.query is interpolated into the user message. If your application is user-facing, sanitise it first — strip control characters, enforce a max length (e.g. 1 000 characters), and never place it in the system prompt where it could override your grounding instructions.Before you ship or iterate, verify the pipeline with a small set of known-answer queries — called a golden set. A golden set has three fields per example: question, expected_answer (or key facts that must appear), and expected_source (the filename the answer should come from).
# smoke_test.py
from generate import generate
GOLDEN_SET = [
{
"question": "What is the document retention policy for financial records?",
"must_contain": ["7 years", "financial"],
"expected_source": "policies.pdf",
},
{
"question": "Which encryption standard is required for data at rest?",
"must_contain": ["AES-256"],
"expected_source": "security-guide.pdf",
},
]
def run_smoke_tests() -> None:
passed = 0
for i, test in enumerate(GOLDEN_SET, start=1):
result = generate(test["question"])
answer = result["answer"].lower()
source_files = [s["filename"] for s in result["sources"]]
content_ok = all(kw.lower() in answer for kw in test["must_contain"])
source_ok = any(
test["expected_source"] in sf for sf in source_files
)
status = "PASS" if (content_ok and source_ok) else "FAIL"
if status == "PASS":
passed += 1
print(f"\nTest {i}: {status}")
print(f" Question: {test['question']}")
if not content_ok:
missing = [kw for kw in test["must_contain"] if kw.lower() not in answer]
print(f" MISSING keywords: {missing}")
if not source_ok:
print(f" WRONG source. Got: {source_files}, expected: {test['expected_source']}")
print(f"\n{passed}/{len(GOLDEN_SET)} tests passed.")
if __name__ == "__main__":
run_smoke_tests()
Common failure patterns and their fixes:
| Failure symptom | Likely cause | Fix |
|---|---|---|
| Wrong source retrieved | Chunk too large, correct fact diluted | Reduce CHUNK_SIZE |
| Answer ignores context | Model hallucinating despite system prompt | Strengthen system prompt constraints |
| No results (score = 0) | Query model != document model | Check EMBEDDING_MODEL constant |
| Keywords present but answer wrong | Retrieved chunks lack the answer | Check the raw source document; may need re-OCR |
| All scores < 0.3 | Corpus too small or off-domain | Add more relevant documents or switch embedding model |
For more rigorous evaluation — faithfulness scores, precision@k, MRR — see Lesson 7: Evaluation & Quality, which introduces RAGAS and LLM-as-judge evaluation.
Questions & Answers
metadata={"hnsw:space": "cosine"} — the default is L2, which produces different distance scales. Second, verify that the query embedding and document embeddings use exactly the same model name string. Third, inspect your cleaning step: if chunks contain mostly boilerplate (headers, page numbers, repeated disclaimers), the embeddings encode noise rather than meaning. Try printing a random sample of 10 chunks before embedding to check their quality.collection.delete(where={"source": old_path}) and re-embed the new version. Production approaches and full re-indexing strategies are covered in Lesson 9: Production RAG.Document, Node, VectorStoreIndex, and RetrieverQueryEngine. Switching to a framework is straightforward once you understand the underlying mechanics.Key Takeaways
- The pipeline has five distinct stages: ingest, chunk, embed, store, retrieve + generate. Keep them modular — you will swap the embedding model, vector store, or LLM without touching the others.
- Metadata is load-bearing: every chunk must carry
source,filename, and position. Lose the metadata and you cannot cite sources, which is the primary value proposition of RAG over a vanilla LLM. - Same model, always: use exactly the same
EMBEDDING_MODELconstant for both ingestion and query-time embedding. A mismatch is silent and produces near-random retrieval. - Cosine space for sentence embeddings: build your Chroma collection with
"hnsw:space": "cosine"— sentence embedding models optimise for directional similarity, not L2 distance. - System prompt constraints drive faithfulness: instruct the model explicitly to cite sources and refuse to go beyond the provided context. This is what separates a grounded RAG answer from a hallucination.
- Golden-set smoke tests catch regressions early: even three well-chosen known-answer queries will surface the most common failure modes before they reach users.
Next Steps: Lesson 6: Advanced Retrieval