Building a Basic RAG Pipeline

60 min intermediate Lesson 5

Learning Outcomes

  • Ingest documents from PDFs and plain text into a normalised Python structure
  • Chunk and embed documents using langchain-text-splitters and sentence-transformers
  • Store and query embeddings in a local Chroma vector database
  • Assemble retrieved context into a grounded prompt and generate a cited answer with the Anthropic SDK
  • Verify the pipeline end-to-end against a set of known-answer test queries

Lesson Plan

Segment Duration Topic
Intro 3 min Pipeline overview and what we're building
Step 1–2 10 min Ingestion and text cleaning
Step 3–4 12 min Chunking and embedding
Step 5 10 min Vector storage with Chroma
Step 6 10 min Retrieval and context assembly
Step 7–8 10 min Grounded generation and smoke-testing
Wrap-up 5 min Key takeaways and what's next

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10 or later
  • An Anthropic API key set as the ANTHROPIC_API_KEY environment variable
  • The following packages installed (one pip install shown in Step 1): pypdf, langchain-text-splitters, sentence-transformers, chromadb, anthropic
  • A small corpus of documents to index — two or three PDFs or .txt files work fine for this lesson

1 Project Layout and Dependency Install

Start from a clean project folder. The layout below keeps pipeline stages clearly separated so you can swap components later without touching unrelated code:

rag-pipeline/
├── data/           # raw source documents (PDFs, .txt, .md)
├── ingest.py       # load + clean + chunk + embed + store
├── retrieve.py     # embed query, similarity search, assemble context
├── generate.py     # call the LLM with grounded prompt
├── pipeline.py     # end-to-end runner
└── smoke_test.py   # verify the pipeline against known queries

Install all runtime dependencies in one command:

pip install pypdf langchain-text-splitters sentence-transformers chromadb anthropic

Verified package names as of 2026-05-31:

  • pypdf — pure-Python PDF reader
  • langchain-text-splitters — standalone text splitting utilities from LangChain
  • sentence-transformers — embedding and reranker models from Hugging Face
  • chromadb — open-source vector database with a Python-native API
  • anthropic — official Anthropic Python SDK

Pin these to a requirements.txt for reproducibility:

pip freeze | grep -E "pypdf|langchain|sentence|chromadb|anthropic" > requirements.txt
TIP
Virtual Environments
Always work in a virtual environment: python -m venv .venv && source .venv/bin/activate. The sentence-transformers wheel downloads model weights on first use, so allow a few minutes and a few hundred MB of disk space the first time.

2 Document Ingestion and Text Cleaning

The first stage converts raw files into normalised Python dicts. Each dict carries the text content plus metadata fields you will later store alongside embeddings in Chroma.

# ingest.py  -- Part 1: load raw documents
import re
from pathlib import Path
from pypdf import PdfReader


def load_pdf(path: Path) -> dict:
    """Extract text from every page of a PDF and return a doc dict."""
    reader = PdfReader(str(path))
    pages = []
    for page_num, page in enumerate(reader.pages):
        raw = page.extract_text() or ""
        pages.append({"page": page_num + 1, "text": raw})
    full_text = "\n\n".join(p["text"] for p in pages)
    return {
        "source": str(path),
        "filename": path.name,
        "type": "pdf",
        "text": full_text,
        "pages": len(reader.pages),
    }


def load_text(path: Path) -> dict:
    """Load a plain-text or Markdown file."""
    text = path.read_text(encoding="utf-8")
    return {
        "source": str(path),
        "filename": path.name,
        "type": "text",
        "text": text,
        "pages": 1,
    }


def load_corpus(data_dir: str = "data") -> list[dict]:
    """Load all supported files from a directory."""
    corpus = []
    for path in sorted(Path(data_dir).iterdir()):
        if path.suffix.lower() == ".pdf":
            corpus.append(load_pdf(path))
        elif path.suffix.lower() in {".txt", ".md"}:
            corpus.append(load_text(path))
    return corpus

After loading, run a lightweight cleaning pass to remove noise that degrades embedding quality:

# ingest.py  -- Part 2: text cleaning
def clean_text(text: str) -> str:
    """Remove common PDF artefacts and normalise whitespace."""
    # Collapse runs of whitespace and blank lines
    text = re.sub(r"[ \t]+", " ", text)
    text = re.sub(r"\n{3,}", "\n\n", text)
    # Remove page-number lines (lines that are only a digit)
    text = re.sub(r"(?m)^\s*\d+\s*$", "", text)
    # Remove repeated header/footer patterns (heuristic: same line 3+ times)
    lines = text.split("\n")
    from collections import Counter
    line_counts = Counter(ln.strip() for ln in lines if ln.strip())
    boilerplate = {ln for ln, cnt in line_counts.items() if cnt >= 3 and len(ln) < 120}
    cleaned = [ln for ln in lines if ln.strip() not in boilerplate]
    return "\n".join(cleaned).strip()

The cleaning heuristics above handle the most common issues in digitised PDFs: duplicate headers, page numbers, and excessive whitespace. For scanned PDFs you will need an OCR layer (pdfplumber or Tesseract) — that is outside scope here.

WARNING
PDF Text Quality
pypdf extracts the text layer of PDFs. Scanned documents produce empty or garbled text because they are images, not text. Check page.extract_text() output before indexing — if you get back empty strings your document needs OCR first.
Source type Load function Notes
PDF (digital) load_pdf() Uses pypdf PdfReader
Plain text / Markdown load_text() UTF-8 read
Web page requests.get() + BeautifulSoup Not shown here; same dict shape
NOTE
Metadata Matters
Store source, filename, and page (or section) on every chunk. You cannot cite a source at generation time if you haven't tracked it through the pipeline. Good metadata is what separates a demo from a production system.

3 Chunking with RecursiveCharacterTextSplitter

Chunking converts a long document text into retrieval units. The Lesson 4 deep-dive covers all major strategies; here we use RecursiveCharacterTextSplitter — the best default for prose documents.

It tries to split on \n\n first (paragraph boundary), then \n (line), then space, so chunk boundaries naturally align with sentence and paragraph edges rather than cutting mid-sentence.

# ingest.py  -- Part 3: chunking
from langchain_text_splitters import RecursiveCharacterTextSplitter


CHUNK_SIZE = 512        # characters, not tokens
CHUNK_OVERLAP = 64      # carry the last 64 chars into the next chunk


def chunk_document(doc: dict) -> list[dict]:
    """Split a document into overlapping text chunks, preserving metadata."""
    splitter = RecursiveCharacterTextSplitter(
        chunk_size=CHUNK_SIZE,
        chunk_overlap=CHUNK_OVERLAP,
        separators=["\n\n", "\n", ". ", " ", ""],
    )
    texts = splitter.split_text(clean_text(doc["text"]))
    chunks = []
    for i, text in enumerate(texts):
        chunks.append({
            "id": f"{doc['filename']}-chunk-{i:04d}",
            "text": text,
            "source": doc["source"],
            "filename": doc["filename"],
            "chunk_index": i,
            "total_chunks": len(texts),
        })
    return chunks


def chunk_corpus(docs: list[dict]) -> list[dict]:
    all_chunks = []
    for doc in docs:
        all_chunks.extend(chunk_document(doc))
    print(f"Chunked {len(docs)} documents into {len(all_chunks)} chunks.")
    return all_chunks

Chunk size trade-offs are worth internalising before you tune:

Chunk size Retrieval precision Context richness Risk
Small (128–256 chars) High — narrow match Low — loses surrounding context Answer lacks detail
Medium (512–800 chars) Balanced Balanced Best starting point
Large (1 200+ chars) Low — too much noise High — rich context Irrelevant content dilutes the prompt

The 64-character overlap carries sentence fragments across boundaries, preventing the retriever from missing a fact that straddles a chunk edge.

TIP
Tokens vs Characters
LLM context limits are measured in tokens, not characters. A rough rule: 512 characters ≈ 100–150 tokens for English prose, depending on vocabulary. At 512-char chunks you can safely fit 5–10 retrieved chunks in the context window alongside the user question and system instructions.

4 Generating Embeddings with sentence-transformers

An embedding is a dense vector that encodes the meaning of a piece of text. Mathematically similar texts produce vectors that are close together (high cosine similarity). This is how the retriever finds relevant chunks without keyword matching.

We use sentence-transformers with the all-MiniLM-L6-v2 model — 384-dimensional vectors, fast CPU inference, strong English-language retrieval, and no API call required.

# ingest.py  -- Part 4: embedding
from sentence_transformers import SentenceTransformer

EMBEDDING_MODEL = "sentence-transformers/all-MiniLM-L6-v2"

# Load once; model weights are cached after the first download
_model: SentenceTransformer | None = None


def get_model() -> SentenceTransformer:
    global _model
    if _model is None:
        _model = SentenceTransformer(EMBEDDING_MODEL)
    return _model


def embed_chunks(chunks: list[dict]) -> list[dict]:
    """Add an 'embedding' key (list of floats) to each chunk dict."""
    model = get_model()
    texts = [c["text"] for c in chunks]
    # encode() returns a numpy array of shape (N, 384)
    vectors = model.encode(texts, batch_size=64, show_progress_bar=True)
    for chunk, vector in zip(chunks, vectors):
        chunk["embedding"] = vector.tolist()   # Chroma expects a plain list
    return chunks

A few things to note about this implementation:

  • batch_size=64 processes 64 chunks per forward pass, which is efficient on CPU. Raise to 128 or 256 if you have a GPU.
  • encode() returns numpy.ndarray; Chroma's Python client expects list[float], so we call .tolist().
  • The model is loaded once at import time and reused — never reinstantiate per chunk or per query.

Embedding the query at retrieval time must use the same model. Using different models for documents and queries is one of the most common RAG bugs and produces near-random retrieval quality.

WARNING
Model Consistency
If you switch embedding models (e.g. upgrading to a higher-dimensional model), you must re-embed and re-index your entire corpus. Mixing old and new embeddings in the same vector store breaks similarity search — the vectors live in different geometric spaces.
NOTE
Choosing an Embedding Model
For English-only corpora, all-MiniLM-L6-v2 is a good default. For multilingual content, paraphrase-multilingual-MiniLM-L12-v2 covers 50+ languages. For the highest-quality English retrieval, check the MTEB leaderboard for current top models — it is updated continuously.

5 Storing Embeddings in Chroma

Chroma is an open-source vector database with a Python-native API and no infrastructure dependencies for local use. PersistentClient writes to disk so your index survives restarts without re-embedding.

# ingest.py  -- Part 5: store in Chroma
import chromadb

CHROMA_PATH = "./chroma_db"
COLLECTION_NAME = "rag_docs"


def get_collection(reset: bool = False) -> chromadb.Collection:
    client = chromadb.PersistentClient(path=CHROMA_PATH)
    if reset:
        try:
            client.delete_collection(COLLECTION_NAME)
        except Exception:
            pass
    collection = client.get_or_create_collection(
        name=COLLECTION_NAME,
        metadata={"hnsw:space": "cosine"},   # use cosine distance
    )
    return collection


def store_chunks(chunks: list[dict], reset: bool = False) -> None:
    """Upsert all chunks (with embeddings) into the Chroma collection."""
    collection = get_collection(reset=reset)

    ids = [c["id"] for c in chunks]
    embeddings = [c["embedding"] for c in chunks]
    documents = [c["text"] for c in chunks]
    metadatas = [
        {
            "source": c["source"],
            "filename": c["filename"],
            "chunk_index": c["chunk_index"],
        }
        for c in chunks
    ]

    # Chroma's add() is idempotent when IDs already exist (upsert behaviour)
    collection.add(
        ids=ids,
        embeddings=embeddings,
        documents=documents,
        metadatas=metadatas,
    )
    print(f"Stored {len(chunks)} chunks. Collection count: {collection.count()}")

The "hnsw:space": "cosine" metadata key tells Chroma to build the HNSW index using cosine distance rather than L2 (Euclidean). Cosine distance is the right choice for sentence embeddings because it measures directional similarity, which is what the embedding model was trained to optimise.

Chroma client Storage When to use
EphemeralClient() RAM only Unit tests, CI
PersistentClient(path=...) Local disk Development, single-machine production
HttpClient(host=..., port=...) Remote server Multi-process, containerised

Now wire up the full ingestion pipeline:

# ingest.py  -- Part 6: main ingestion runner
def run_ingestion(data_dir: str = "data", reset: bool = False) -> None:
    docs = load_corpus(data_dir)
    chunks = chunk_corpus(docs)
    chunks = embed_chunks(chunks)
    store_chunks(chunks, reset=reset)
    print("Ingestion complete.")


if __name__ == "__main__":
    run_ingestion(reset=True)

Run it:

python ingest.py

You should see progress output from encode() and a final count. If you rerun without reset=True, add() is idempotent — existing IDs are updated, new IDs are inserted.

TIP
Incremental Updates
In production you will add new documents without rebuilding from scratch. Use get_or_create_collection() without reset=True and call collection.add() with only the new chunks. Track which files are already indexed (e.g. by storing a hash in a sidecar SQLite DB) to avoid re-embedding unchanged files.

6 Retrieval and Context Assembly

Retrieval is the core of RAG: embed the user's question, find the nearest document chunks, and assemble them into a context block the LLM can reason over.

# retrieve.py
import chromadb
from sentence_transformers import SentenceTransformer

from ingest import CHROMA_PATH, COLLECTION_NAME, EMBEDDING_MODEL

TOP_K = 5   # number of chunks to retrieve


def retrieve(query: str, top_k: int = TOP_K) -> list[dict]:
    """Embed query, search Chroma, return ranked results."""
    model = SentenceTransformer(EMBEDDING_MODEL)
    query_vector = model.encode(query).tolist()

    client = chromadb.PersistentClient(path=CHROMA_PATH)
    collection = client.get_or_create_collection(
        name=COLLECTION_NAME,
        metadata={"hnsw:space": "cosine"},
    )
    results = collection.query(
        query_embeddings=[query_vector],
        n_results=top_k,
        include=["documents", "metadatas", "distances"],
    )

    hits = []
    for doc, meta, dist in zip(
        results["documents"][0],
        results["metadatas"][0],
        results["distances"][0],
    ):
        hits.append({
            "text": doc,
            "source": meta.get("source", "unknown"),
            "filename": meta.get("filename", "unknown"),
            "chunk_index": meta.get("chunk_index", -1),
            "score": 1.0 - dist,   # convert distance to similarity (0..1)
        })
    return hits


def assemble_context(hits: list[dict]) -> str:
    """Format retrieved chunks into a numbered context block."""
    parts = []
    for i, hit in enumerate(hits, start=1):
        parts.append(
            f"[Source {i}: {hit['filename']}, chunk {hit['chunk_index']}]\n"
            f"{hit['text']}"
        )
    return "\n\n---\n\n".join(parts)

The score field (1 - cosine distance) gives you a relevance signal in [0, 1]. You can use it to filter low-quality hits before assembly:

MIN_SCORE = 0.30

def retrieve_filtered(query: str, top_k: int = TOP_K) -> list[dict]:
    hits = retrieve(query, top_k=top_k * 2)   # over-fetch, then filter
    return [h for h in hits if h["score"] >= MIN_SCORE][:top_k]

Over-fetching then filtering is a practical pattern for skipping the noise chunks that fall below the relevance threshold without tuning n_results per query.

NOTE
What top_k Should I Use?
Start with 5. Too few (1–2) and you risk missing the relevant chunk; too many (10+) and the context window fills with noise that confuses the LLM. In Lesson 6 you will add a cross-encoder re-ranker that lets you retrieve 20 candidates and re-score them for precision.
WARNING
Same Model for Query and Documents
The retrieve() function loads the same EMBEDDING_MODEL constant defined in ingest.py. Never hardcode the model name in two places — always import from a single source-of-truth constant. A mismatch here is invisible at startup but produces retrieval scores near 0.

7 Grounded Generation with Source Citation

The final stage passes the assembled context to Claude alongside the user's question. The system prompt instructs Claude to ground its answer strictly in the retrieved context and to cite the source labels introduced in assemble_context().

# generate.py
import os
import anthropic

from retrieve import retrieve_filtered, assemble_context

SYSTEM_PROMPT = """You are a precise research assistant. Answer the user's question
using ONLY the information in the provided context sources. Rules:

1. Base every factual claim on a specific source. Cite it inline as [Source N].
2. If the context does not contain enough information to answer, say so explicitly.
   Do NOT fabricate facts or draw on general knowledge.
3. Be concise. Prefer one clear paragraph over a verbose list unless a list is
   genuinely clearer.
4. If two sources conflict, acknowledge the conflict and cite both."""

GENERATION_MODEL = "claude-sonnet-4-6"


def generate(query: str, top_k: int = 5) -> dict:
    """Full RAG cycle: retrieve -> assemble -> generate."""
    hits = retrieve_filtered(query, top_k=top_k)

    if not hits:
        return {
            "answer": "No relevant documents found for this query.",
            "sources": [],
            "hits": [],
        }

    context = assemble_context(hits)
    user_message = f"""Context sources:

{context}

---

Question: {query}

Answer (cite sources inline):"""

    client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
    response = client.messages.create(
        model=GENERATION_MODEL,
        max_tokens=1024,
        system=SYSTEM_PROMPT,
        messages=[{"role": "user", "content": user_message}],
    )

    answer = response.content[0].text
    sources = [
        {"label": f"Source {i+1}", "filename": h["filename"], "score": h["score"]}
        for i, h in enumerate(hits)
    ]
    return {"answer": answer, "sources": sources, "hits": hits}

Note the structure of the prompt:

  • The system prompt sets hard constraints (cite sources, no fabrication, acknowledge gaps).
  • The user message embeds the numbered context block then asks the question below it. Separating context from question with a horizontal rule gives the model a clear boundary.
  • The cited [Source N] labels map directly to the filenames in sources, so you can present them in a UI.

Wire everything into pipeline.py:

# pipeline.py
from generate import generate
import json

if __name__ == "__main__":
    query = input("Question: ").strip()
    result = generate(query)
    print("\n=== Answer ===")
    print(result["answer"])
    print("\n=== Sources used ===")
    for s in result["sources"]:
        print(f"  {s['label']}: {s['filename']} (relevance {s['score']:.2f})")

Run it:

python pipeline.py
TIP
Streaming Responses
For long answers, pass stream=True to client.messages.create() and iterate over the stream to print tokens as they arrive. The Anthropic SDK supports streaming natively — see the SDK docs for the stream() context manager pattern.
WARNING
Never Prompt-Inject User Input
The user's query is interpolated into the user message. If your application is user-facing, sanitise it first — strip control characters, enforce a max length (e.g. 1 000 characters), and never place it in the system prompt where it could override your grounding instructions.

8 Smoke-Testing the Pipeline

Before you ship or iterate, verify the pipeline with a small set of known-answer queries — called a golden set. A golden set has three fields per example: question, expected_answer (or key facts that must appear), and expected_source (the filename the answer should come from).

# smoke_test.py
from generate import generate

GOLDEN_SET = [
    {
        "question": "What is the document retention policy for financial records?",
        "must_contain": ["7 years", "financial"],
        "expected_source": "policies.pdf",
    },
    {
        "question": "Which encryption standard is required for data at rest?",
        "must_contain": ["AES-256"],
        "expected_source": "security-guide.pdf",
    },
]


def run_smoke_tests() -> None:
    passed = 0
    for i, test in enumerate(GOLDEN_SET, start=1):
        result = generate(test["question"])
        answer = result["answer"].lower()
        source_files = [s["filename"] for s in result["sources"]]

        content_ok = all(kw.lower() in answer for kw in test["must_contain"])
        source_ok = any(
            test["expected_source"] in sf for sf in source_files
        )

        status = "PASS" if (content_ok and source_ok) else "FAIL"
        if status == "PASS":
            passed += 1

        print(f"\nTest {i}: {status}")
        print(f"  Question: {test['question']}")
        if not content_ok:
            missing = [kw for kw in test["must_contain"] if kw.lower() not in answer]
            print(f"  MISSING keywords: {missing}")
        if not source_ok:
            print(f"  WRONG source. Got: {source_files}, expected: {test['expected_source']}")

    print(f"\n{passed}/{len(GOLDEN_SET)} tests passed.")


if __name__ == "__main__":
    run_smoke_tests()

Common failure patterns and their fixes:

Failure symptom Likely cause Fix
Wrong source retrieved Chunk too large, correct fact diluted Reduce CHUNK_SIZE
Answer ignores context Model hallucinating despite system prompt Strengthen system prompt constraints
No results (score = 0) Query model != document model Check EMBEDDING_MODEL constant
Keywords present but answer wrong Retrieved chunks lack the answer Check the raw source document; may need re-OCR
All scores < 0.3 Corpus too small or off-domain Add more relevant documents or switch embedding model

For more rigorous evaluation — faithfulness scores, precision@k, MRR — see Lesson 7: Evaluation & Quality, which introduces RAGAS and LLM-as-judge evaluation.

NOTE
Minimum Viable Test Set
You do not need 100 test queries to start. Three to five carefully chosen questions — one per major document type in your corpus — catch the most common failure modes quickly. Expand the golden set incrementally as you encounter real user queries that the pipeline gets wrong.

Questions & Answers

Q: My retrieval scores are consistently below 0.25 even for queries I know are in the corpus. What's wrong?
Three likely causes. First, check that you built the Chroma collection with metadata={"hnsw:space": "cosine"} — the default is L2, which produces different distance scales. Second, verify that the query embedding and document embeddings use exactly the same model name string. Third, inspect your cleaning step: if chunks contain mostly boilerplate (headers, page numbers, repeated disclaimers), the embeddings encode noise rather than meaning. Try printing a random sample of 10 chunks before embedding to check their quality.
Q: The LLM keeps adding facts that aren't in the retrieved context. How do I stop it?
This is faithfulness failure — the model is drawing on parametric (training) knowledge instead of restricting to the context. First, harden the system prompt: add "If a fact is not explicitly stated in the sources above, do not state it." Second, consider using a lower temperature (0.0–0.2) for factual Q&A. Third, add a faithfulness evaluation metric in your testing (see Lesson 7) so you can measure and track improvement. Some hallucination is irreducible with pure prompting; for high-stakes use cases, add a post-generation verification step.
Q: How do I handle documents that are updated regularly — like internal policies that change quarterly?
Build an incremental update flow. Assign each document a hash (e.g. SHA-256 of its bytes) and store it alongside the Chroma collection. On each ingestion run, hash the incoming file: if the hash matches what's stored, skip it; if it differs, delete the old chunks by calling collection.delete(where={"source": old_path}) and re-embed the new version. Production approaches and full re-indexing strategies are covered in Lesson 9: Production RAG.
Q: Should I use LlamaIndex or LangChain instead of building this from scratch?
For a learning exercise, building from scratch is far more instructive — you see exactly what each stage does and can debug at any layer. In production, LlamaIndex and LangChain both provide well-tested, modular implementations of every component in this lesson, plus integrations with dozens of vector stores, document loaders, and LLM providers. The architecture you've built here maps directly onto LlamaIndex concepts: Document, Node, VectorStoreIndex, and RetrieverQueryEngine. Switching to a framework is straightforward once you understand the underlying mechanics.
Q: The pipeline works in my tests but gives poor answers on real user queries. The test queries pass but production fails. Why?
Your golden set queries are likely too similar to the training data or too easy. Real user queries are noisier, more ambiguous, and often use different terminology than the source documents (vocabulary mismatch). Three mitigations: (1) Collect real failed queries and add them to the golden set. (2) Add query expansion in retrieval — generate 2–3 alternative phrasings of the user question and retrieve for all of them, then deduplicate (covered in Lesson 6). (3) Log every query and its top-3 retrieved chunk scores so you can identify systematic misses. You cannot fix what you cannot observe.

Key Takeaways

  1. The pipeline has five distinct stages: ingest, chunk, embed, store, retrieve + generate. Keep them modular — you will swap the embedding model, vector store, or LLM without touching the others.
  2. Metadata is load-bearing: every chunk must carry source, filename, and position. Lose the metadata and you cannot cite sources, which is the primary value proposition of RAG over a vanilla LLM.
  3. Same model, always: use exactly the same EMBEDDING_MODEL constant for both ingestion and query-time embedding. A mismatch is silent and produces near-random retrieval.
  4. Cosine space for sentence embeddings: build your Chroma collection with "hnsw:space": "cosine" — sentence embedding models optimise for directional similarity, not L2 distance.
  5. System prompt constraints drive faithfulness: instruct the model explicitly to cite sources and refuse to go beyond the provided context. This is what separates a grounded RAG answer from a hallucination.
  6. Golden-set smoke tests catch regressions early: even three well-chosen known-answer queries will surface the most common failure modes before they reach users.

Next Steps: Lesson 6: Advanced Retrieval