RAG Principles
The Core RAG Contract
A RAG system makes one promise: answers are grounded in your documents, not invented from model weights. Every principle below exists to honour that promise, or to make it sustainable at production scale.
Principle 1 — Retrieve Before You Generate
Always retrieve context before calling the generation model. The LLM's role is synthesis and formatting, not recall.
| Anti-pattern | Correct pattern |
|---|---|
| Ask LLM to recall facts from training | Retrieve relevant chunks, pass as context |
| Let the model hedge ("I'm not sure…") | Pin answers to retrieved text; if nothing retrieves, say so |
| One retrieval path for all query types | Route exact-match queries (IDs, dates) to SQL; semantic queries to the vector store |
def answer(query, retriever, llm):
chunks = retriever.retrieve(query, top_k=5)
if not chunks:
return "No relevant documents found."
context = "\n\n".join(c.text for c in chunks)
return llm.complete(f"Answer using only the context below.\n\n{context}\n\nQuestion: {query}")
Principle 2 — Chunk for the Question, Not the Document
The chunk is the unit of retrieval. Design chunks around the questions users will ask, not the structure of the source document.
| Strategy | Best for | Watch out for |
|---|---|---|
| Fixed-size (512 tokens, 64-token overlap) | Quick prototyping | Splits mid-sentence, mid-table |
| Recursive character split | Prose documents | Ignores semantic boundaries |
| Structure-aware (headings, sections) | Markdown, HTML | Needs a parser per format |
| Parent-child (small retrieval, large context) | Precision + full context | More complex index design |
Enrich every chunk with metadata (source path, section heading, page number, doc date) — free signal for filtering and citation.
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=64)
chunks = splitter.split_text(document_text)
docs = [{"text": c, "metadata": {"source": path, "section": heading}} for c in chunks]
Principle 3 — Measure Retrieval Separately from Generation
End-to-end answer quality conflates two failure modes: retrieval miss and generation hallucination. Measure both independently.
Retrieval metrics
| Metric | What it measures |
|---|---|
| Precision@k | Fraction of top-k results that are relevant |
| Recall@k | Fraction of all relevant docs that appear in top-k |
| MRR | How high the first relevant result ranks |
| Hit rate | Whether any relevant chunk appears in top-k (binary) |
Generation metrics
| Metric | What it measures | Tool |
|---|---|---|
| Faithfulness | Claims supported by retrieved context | RAGAS, LLM-as-judge |
| Answer relevance | Does the answer address the question? | RAGAS, human eval |
| Answer correctness | Match against known ground truth | Semantic similarity or exact match |
Build a golden test set early: fifty hand-labelled question/answer/source triples reveal more than thousands of unlabelled queries.
Principle 4 — Ground Answers and Cite Sources
Every factual claim should map back to a retrievable chunk. Citations prevent the model blending training-data knowledge with retrieved context and give you a free audit trail.
Answer using ONLY the numbered sources. After each claim add [N].
If sources are insufficient, say "Insufficient context."
[1] {chunk_1_text}
[2] {chunk_2_text}
Question: {query}
Principle 5 — Use Hybrid Search; Re-rank Before Generation
Pure vector search fails on exact terms (product codes, version strings). Pure BM25 misses paraphrase. Combine both via Reciprocal Rank Fusion, then re-rank.
Pipeline: vector search (top 20–50) → BM25 merge (RRF) → cross-encoder re-rank (top 5) → LLM generation.
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")
def rerank(query, chunks, top_n=5):
scores = reranker.predict([(query, c) for c in chunks])
return [c for _, c in sorted(zip(scores, chunks), reverse=True)[:top_n]]
Principle 6 — Cache Aggressively at Every Layer
| Cache layer | Key | TTL |
|---|---|---|
| Query embedding | hash(query_text) |
Days–weeks |
| Retrieval results | hash(query_embedding) |
Hours–days |
| Generated answer | hash(query + chunk_ids) |
Minutes–hours |
| Chunk embedding | hash(chunk_text) |
Permanent until chunk changes |
Couple TTL to your document update cadence. Cached answers become stale when source documents change.
Principle 7 — Plan for Document Updates from Day One
| Strategy | Use when |
|---|---|
| Full re-index | Small corpus, infrequent updates |
| Incremental upsert (hash each chunk) | Medium corpus, regular updates |
Append-only with valid_until metadata |
Point-in-time queries needed |
| Event-driven (webhook triggers re-index) | Large corpus, real-time accuracy |
Track deletions explicitly. Orphaned chunks produce confident-sounding answers from documents that no longer exist.
import hashlib
def upsert_document(doc_path, new_text, vector_store, embedder, chunk_fn):
doc_hash = hashlib.sha256(new_text.encode()).hexdigest()
existing = vector_store.get(where={"source": doc_path})
if existing and existing["metadata"][0].get("hash") == doc_hash:
return # unchanged
vector_store.delete(where={"source": doc_path})
for i, chunk in enumerate(chunk_fn(new_text)):
vector_store.add(
ids=[f"{doc_path}::{i}"],
embeddings=[embedder.encode(chunk).tolist()],
documents=[chunk],
metadatas=[{"source": doc_path, "hash": doc_hash}],
)
Principle 8 — Match the Embedding Model to Your Domain
Consistency rule: index-time and query-time embeddings must use the same model and settings. Mixing models silently destroys retrieval quality.
| Model family | Strengths |
|---|---|
OpenAI text-embedding-3-* |
High general quality, managed |
BAAI/bge-large-en-v1.5 |
Strong open-source baseline |
intfloat/e5-large-v2 |
Good with instruction prefixes |
nomic-embed-text |
Long context (8k tokens), Apache |
| Domain fine-tuned | Best on specialised corpora |
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-large-en-v1.5")
query_vec = model.encode("query: What is retrieval-augmented generation?")
Principle 9 — Fail Gracefully When Retrieval Comes Up Empty
Return a clear "no result" response rather than letting the LLM hallucinate from training weights.
MIN_SCORE = 0.70
def safe_answer(query, retriever, llm):
results = retriever.retrieve(query, top_k=5, score_threshold=MIN_SCORE)
if not results:
return {"answer": "No relevant information found.", "sources": []}
context = "\n\n".join(r.text for r in results)
return {"answer": llm.complete(build_prompt(query, context)),
"sources": [r.metadata["source"] for r in results]}
Tune MIN_SCORE iteratively: too high and no-result responses become common; too low and hallucinations return.