RAG Principles

Reference intermediate

The Core RAG Contract

A RAG system makes one promise: answers are grounded in your documents, not invented from model weights. Every principle below exists to honour that promise, or to make it sustainable at production scale.


Principle 1 — Retrieve Before You Generate

Always retrieve context before calling the generation model. The LLM's role is synthesis and formatting, not recall.

Anti-pattern Correct pattern
Ask LLM to recall facts from training Retrieve relevant chunks, pass as context
Let the model hedge ("I'm not sure…") Pin answers to retrieved text; if nothing retrieves, say so
One retrieval path for all query types Route exact-match queries (IDs, dates) to SQL; semantic queries to the vector store
def answer(query, retriever, llm):
    chunks = retriever.retrieve(query, top_k=5)
    if not chunks:
        return "No relevant documents found."
    context = "\n\n".join(c.text for c in chunks)
    return llm.complete(f"Answer using only the context below.\n\n{context}\n\nQuestion: {query}")

Principle 2 — Chunk for the Question, Not the Document

The chunk is the unit of retrieval. Design chunks around the questions users will ask, not the structure of the source document.

Strategy Best for Watch out for
Fixed-size (512 tokens, 64-token overlap) Quick prototyping Splits mid-sentence, mid-table
Recursive character split Prose documents Ignores semantic boundaries
Structure-aware (headings, sections) Markdown, HTML Needs a parser per format
Parent-child (small retrieval, large context) Precision + full context More complex index design

Enrich every chunk with metadata (source path, section heading, page number, doc date) — free signal for filtering and citation.

from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=64)
chunks = splitter.split_text(document_text)
docs = [{"text": c, "metadata": {"source": path, "section": heading}} for c in chunks]

Principle 3 — Measure Retrieval Separately from Generation

End-to-end answer quality conflates two failure modes: retrieval miss and generation hallucination. Measure both independently.

Retrieval metrics

Metric What it measures
Precision@k Fraction of top-k results that are relevant
Recall@k Fraction of all relevant docs that appear in top-k
MRR How high the first relevant result ranks
Hit rate Whether any relevant chunk appears in top-k (binary)

Generation metrics

Metric What it measures Tool
Faithfulness Claims supported by retrieved context RAGAS, LLM-as-judge
Answer relevance Does the answer address the question? RAGAS, human eval
Answer correctness Match against known ground truth Semantic similarity or exact match

Build a golden test set early: fifty hand-labelled question/answer/source triples reveal more than thousands of unlabelled queries.


Principle 4 — Ground Answers and Cite Sources

Every factual claim should map back to a retrievable chunk. Citations prevent the model blending training-data knowledge with retrieved context and give you a free audit trail.

Answer using ONLY the numbered sources. After each claim add [N].
If sources are insufficient, say "Insufficient context."

[1] {chunk_1_text}
[2] {chunk_2_text}

Question: {query}

Principle 5 — Use Hybrid Search; Re-rank Before Generation

Pure vector search fails on exact terms (product codes, version strings). Pure BM25 misses paraphrase. Combine both via Reciprocal Rank Fusion, then re-rank.

Pipeline: vector search (top 20–50) → BM25 merge (RRF) → cross-encoder re-rank (top 5) → LLM generation.

from sentence_transformers import CrossEncoder

reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")

def rerank(query, chunks, top_n=5):
    scores = reranker.predict([(query, c) for c in chunks])
    return [c for _, c in sorted(zip(scores, chunks), reverse=True)[:top_n]]

Principle 6 — Cache Aggressively at Every Layer

Cache layer Key TTL
Query embedding hash(query_text) Days–weeks
Retrieval results hash(query_embedding) Hours–days
Generated answer hash(query + chunk_ids) Minutes–hours
Chunk embedding hash(chunk_text) Permanent until chunk changes

Couple TTL to your document update cadence. Cached answers become stale when source documents change.


Principle 7 — Plan for Document Updates from Day One

Strategy Use when
Full re-index Small corpus, infrequent updates
Incremental upsert (hash each chunk) Medium corpus, regular updates
Append-only with valid_until metadata Point-in-time queries needed
Event-driven (webhook triggers re-index) Large corpus, real-time accuracy

Track deletions explicitly. Orphaned chunks produce confident-sounding answers from documents that no longer exist.

import hashlib

def upsert_document(doc_path, new_text, vector_store, embedder, chunk_fn):
    doc_hash = hashlib.sha256(new_text.encode()).hexdigest()
    existing = vector_store.get(where={"source": doc_path})
    if existing and existing["metadata"][0].get("hash") == doc_hash:
        return  # unchanged
    vector_store.delete(where={"source": doc_path})
    for i, chunk in enumerate(chunk_fn(new_text)):
        vector_store.add(
            ids=[f"{doc_path}::{i}"],
            embeddings=[embedder.encode(chunk).tolist()],
            documents=[chunk],
            metadatas=[{"source": doc_path, "hash": doc_hash}],
        )

Principle 8 — Match the Embedding Model to Your Domain

Consistency rule: index-time and query-time embeddings must use the same model and settings. Mixing models silently destroys retrieval quality.

Model family Strengths
OpenAI text-embedding-3-* High general quality, managed
BAAI/bge-large-en-v1.5 Strong open-source baseline
intfloat/e5-large-v2 Good with instruction prefixes
nomic-embed-text Long context (8k tokens), Apache
Domain fine-tuned Best on specialised corpora
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-large-en-v1.5")
query_vec = model.encode("query: What is retrieval-augmented generation?")

Principle 9 — Fail Gracefully When Retrieval Comes Up Empty

Return a clear "no result" response rather than letting the LLM hallucinate from training weights.

MIN_SCORE = 0.70

def safe_answer(query, retriever, llm):
    results = retriever.retrieve(query, top_k=5, score_threshold=MIN_SCORE)
    if not results:
        return {"answer": "No relevant information found.", "sources": []}
    context = "\n\n".join(r.text for r in results)
    return {"answer": llm.complete(build_prompt(query, context)),
            "sources": [r.metadata["source"] for r in results]}

Tune MIN_SCORE iteratively: too high and no-result responses become common; too low and hallucinations return.


Related References