RAG Cheat Sheet
Reference
intermediate
The Canonical RAG Pipeline
Two phases — indexing (offline) and querying (online, per request). Both must use the same embedding model.
| Phase | Step | What Happens |
|---|---|---|
| Index | Load | Ingest documents (PDF, HTML, Markdown, DB) |
| Index | Chunk | Split into retrieval units (see Chunking table) |
| Index | Embed | Chunk → dense vector via encoder model |
| Index | Store | Vectors + metadata written to vector store |
| Query | Embed query | User question → same encoder → query vector |
| Query | Retrieve | ANN search returns top-k nearest chunks |
| Query | Re-rank (opt.) | Cross-encoder rescores top-k |
| Query | Generate | LLM answers using retrieved chunks as context |
Chunking Strategy Comparison
| Strategy | Chunk Size | Best For | Risk |
|---|---|---|---|
| Fixed-size | 256–512 t | Quick prototypes | Breaks sentences at boundaries |
| Recursive char split | 512–1024 t, 10–20% overlap | Markdown, code, general text | Size-driven, not meaning-driven |
| Semantic / topic | ~300–800 t variable | Long-form prose, research docs | Slower; threshold needs tuning |
| Document-structure-aware | Matches natural sections | Docs with H1/H2 hierarchy | Requires DOM or AST parsing |
| Parent–child | Child 128–256 t; parent 512–1024 t | Technical docs: precision + context | Two-level index to maintain |
Default: recursive-character-split, 512 tokens + 10% overlap. Switch to semantic or parent–child when precision stalls.
Distance Metrics
| Metric | When to Use | Notes |
|---|---|---|
| Cosine similarity | Most text use cases | Invariant to magnitude; default for sentence-transformers |
| Dot product | Vectors already L2-normalised | Same as cosine on unit vectors; slightly faster |
| Euclidean (L2) | Absolute position matters | Sensitive to magnitude; rarely best for text |
Metric is set at collection-creation time — cannot change without re-indexing.
Embedding Model Quick Reference
| Model | Dim | Licence | Notes |
|---|---|---|---|
text-embedding-3-small / large |
1536 / 3072 | Commercial API | OpenAI; Matryoshka truncation supported |
BAAI/bge-large-en-v1.5 |
1024 | MIT | Top open-source English model; self-hostable |
intfloat/e5-mistral-7b-instruct |
4096 | MIT | Instruction-tuned; MTEB top tier; GPU needed |
all-MiniLM-L6-v2 |
384 | Apache 2 | Fast CPU baseline; lower quality ceiling |
Vector Store Decision Matrix
| Store | Deployment | Hybrid Search | Best Default For |
|---|---|---|---|
| Chroma | Local / embedded | No | Local dev, notebooks, prototypes |
| pgvector | Self-hosted Postgres | Yes (FTS + vector) | Teams already on Postgres; ACID needed |
| Qdrant | Self-hosted / cloud | Yes (sparse + dense) | High-throughput production; on-prem OK |
| Pinecone | Fully managed SaaS | Yes (sparse + dense) | Managed simplicity; zero infra to run |
HNSW dominates — O(log n) queries, tunable via M (connectivity) and ef_construction (build quality).
Minimal Retrieval Pipeline (Python)
# pip install chromadb sentence-transformers openai
import chromadb
from sentence_transformers import SentenceTransformer
from openai import OpenAI
embedder = SentenceTransformer("BAAI/bge-large-en-v1.5")
chroma = chromadb.PersistentClient(path="./chroma_db")
col = chroma.get_or_create_collection("docs", metadata={"hnsw:space": "cosine"})
llm = OpenAI() # reads OPENAI_API_KEY
def index_chunks(chunks: list[dict]) -> None:
texts = [c["text"] for c in chunks]
vecs = embedder.encode(texts, normalize_embeddings=True).tolist()
col.upsert(ids=[c["id"] for c in chunks], embeddings=vecs,
documents=texts, metadatas=[{"source": c["source"]} for c in chunks])
def retrieve(query: str, k: int = 5) -> list[dict]:
qvec = embedder.encode([query], normalize_embeddings=True).tolist()
res = col.query(query_embeddings=qvec, n_results=k)
return [{"text": doc, "source": meta["source"], "distance": dist}
for doc, meta, dist in zip(res["documents"][0], res["metadatas"][0], res["distances"][0])]
def rag_answer(question: str) -> str:
hits = retrieve(question)
context = "\n\n".join(f"[{i+1}] ({h['source']})\n{h['text']}" for i, h in enumerate(hits))
prompt = (
"Answer using only the context below. Cite as [1], [2], etc. "
"Say 'I don't know' if the context is insufficient.\n\n"
f"Context:\n{context}\n\nQuestion: {question}"
)
resp = llm.chat.completions.create(model="gpt-4o-mini", messages=[{"role": "user", "content": prompt}])
return resp.choices[0].message.content
Advanced Retrieval Techniques
| Technique | Summary | When to Add |
|---|---|---|
| Hybrid search (BM25 + dense) | Keyword + vector via Reciprocal Rank Fusion | Exact terms matter (codes, names, jargon) |
| Re-ranking (cross-encoder) | Heavier model rescores (query, chunk) pairs | Good recall but poor precision |
| Query expansion | Generate 3–5 rewrites; union retrieved sets | Short or ambiguous queries |
| HyDE | Embed a hypothetical ideal answer as the query | Abstract or under-specified questions |
| Parent–child retrieval | Index child chunks; return parent window | Small chunks score well but lack context |
| Metadata pre-filter | Filter by date/dept before ANN search | Multi-tenant; time-sensitive data |
# pip install sentence-transformers
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")
def rerank(query: str, hits: list[dict], top_n: int = 3) -> list[dict]:
scores = reranker.predict([(query, h["text"]) for h in hits])
return [h for _, h in sorted(zip(scores, hits), reverse=True)[:top_n]]
Evaluation Metrics
| Metric | What It Measures | Computed As |
|---|---|---|
| Precision@k | Fraction of top-k that are relevant | relevant_in_k / k |
| Recall@k | Fraction of all relevant docs in top-k | relevant_in_k / total |
| MRR | Rank of the first relevant result | mean(1/rank) across queries |
| Faithfulness | Answer stays within retrieved context | LLM-as-judge or RAGAS |
| Answer relevance | Answer addresses the question | RAGAS answer_relevancy |
| Answer correctness | Factually right vs. ground truth | Labelled eval set required |
Build a 50+ item golden test set before optimising — without one, every change is a guess.
Failure Modes & Fixes
| Symptom | Likely Cause | Fix |
|---|---|---|
| Correct docs not retrieved | Chunks too large; query too vague | Reduce chunk size; add query expansion |
| Answer ignores retrieved context | Prompt doesn't enforce grounding | Add "use only the context" instruction |
| Irrelevant chunks in top-k | Embedding model poor fit for domain | Switch model or add re-ranking |
| Answers go stale | Index not updated after doc changes | Incremental re-index on writes |
| Slow query latency | HNSW ef too high, or no GPU |
Tune ef; use int8 quantisation; add cache |
| Multi-hop questions fail | Single-stage retrieval misses chained facts | Add knowledge graph or multi-step retriever |