Production RAG
Learning Outcomes
- Design a multi-layer caching strategy that cuts retrieval latency and LLM spend
- Implement incremental re-indexing so document updates propagate without a full rebuild
- Version your index and trace which document revision produced any given answer
- Build monitoring that surfaces retrieval-quality drift before users notice it
- Diagnose and recover from the failure modes unique to production RAG
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | What changes when you go to production |
| Steps 1-2 | 14 min | Caching at every layer |
| Step 3 | 9 min | Document updates and re-indexing |
| Step 4 | 7 min | Versioning the index |
| Step 5 | 9 min | Monitoring for quality drift |
| Step 6 | 8 min | Scaling and batch ingestion |
| Step 7 | 2 min | Failure modes quick reference |
| Wrap-up | 3 min | Production operations plan |
Before You Begin
Pre-work:
- Complete Lesson 5: Building a Basic RAG Pipeline -- the pipeline extended here
- Complete Lesson 6: Advanced Retrieval -- hybrid search adds latency caching addresses
- Complete Lesson 7: Evaluation & Quality -- its metrics underpin the monitoring in Step 5
- Skim the Agentic AI course if unfamiliar with tool-calling
Shopping List:
- Python 3.10+ with
pip install chromadb qdrant-client sentence-transformers redis hiredis prometheus-client psycopg2-binary - Redis:
docker run -p 6379:6379 redis:7-alpine - An LLM API key (Anthropic, OpenAI, or local)
A prototype RAG pipeline has one job: retrieve documents and generate an answer. Production adds five more.
| Prototype assumption | Production reality |
|---|---|
| Queries are unique | 60-80% repeat within an hour |
| Documents are static | Docs change daily; deletions without notice |
| One model forever | Model changes force migrations |
| Errors are obvious | Quality drifts slowly and silently |
| Low QPS, single node | Hundreds of concurrent requests |
Three layers to build on Lesson 5: Cache (embedding, query results, answers), Index (versioned, model-tagged), Ops (monitoring, alerting, rollback).
RAG has three expensive operations -- embedding (~10-50 ms), vector search (~20-200 ms), LLM generation (~500 ms+) -- each cacheable with separate TTLs.
Embedding cache (1 hr) -- SHA-256 the query, store raw float32 bytes in Redis. Hit: np.frombuffer(r.get(key), dtype=np.float32). Miss: embed, then r.setex(key, 3600, vec.tobytes()).
Query result cache (5 min) -- key on the vector rounded to 3 decimal places; store the serialised hit list as JSON.
Answer cache (30 min) -- only for stable queries; skip for time-relative or user-specific ones.
| Cache layer | Key | Hit rate | TTL |
|---|---|---|---|
| Embedding | SHA-256 of query text | 70-90% | 1 hr |
| Query results | Rounded vector hash | 40-60% | 5 min |
| Answer | Normalised query string | 20-50% | 30 min |
chunk_id to cache keys and delete on update. If too complex, use a short TTL instead.Nightly full rebuilds break down at scale. The production pattern: re-embed only documents whose content changed.
Content-hash gating -- SHA-256 the text; if the hash matches the stored record, skip re-embedding entirely. When 5% of docs change daily, this cuts ingestion cost by 95%.
import hashlib
def update_document(doc_id, text, collection, registry, chunker, embedder):
h = hashlib.sha256(text.encode()).hexdigest()
if doc_id in registry and registry[doc_id]["hash"] == h:
return # unchanged
if doc_id in registry:
collection.delete(ids=registry[doc_id]["chunk_ids"])
chunks = chunker(text)
ids = [f"{doc_id}::{i}" for i in range(len(chunks))]
collection.add(ids=ids, documents=chunks,
embeddings=[embedder(c).tolist() for c in chunks],
metadatas=[{"doc_id": doc_id, "i": i} for i in range(len(chunks))])
registry[doc_id] = {"hash": h, "chunk_ids": ids,
"version": registry.get(doc_id, {}).get("version", 0) + 1}
Deletion handling -- skipping leaves deleted docs in the index forever. Call collection.delete(ids=registry[doc_id]["chunk_ids"]) then del registry[doc_id].
deleted_at timestamp to chunk metadata and filter at query time instead of hard-deleting -- proves what was available at any past date.Versioning answers the question that has no answer without it: "Last Tuesday's answer was wrong -- what documents did my system retrieve?"
What to version:
| Artifact | How |
|---|---|
| Document registry | Append-only SQL event log |
| Embedding model ID | Model ID in every chunk's metadata |
| Chunking config | chunk_size and overlap in chunk metadata |
Minimal PostgreSQL event log:
CREATE TABLE rag_events (
event_id BIGSERIAL PRIMARY KEY, occurred_at TIMESTAMPTZ DEFAULT NOW(),
event_type TEXT NOT NULL, doc_id TEXT, doc_version INT,
content_hash TEXT, model_id TEXT, chunk_count INT
);
CREATE INDEX ON rag_events (doc_id, occurred_at DESC);
One row per update_document call lets you answer: at time T, what version of doc X was indexed and which model produced its vectors?
Embedding model migration -- new model vectors are geometrically incompatible with old ones. Safe procedure: build a parallel collection, shadow-test at 5%, compare precision@5 and MRR, ramp to 100%, retire the old.
Quality drift kills user trust slowly -- new topics appear in questions before documents are indexed, or popular docs are deleted without replacement.
Expose three Prometheus metrics (scrape with Grafana) and add a 1% faithfulness-sampling loop (LLM-as-judge, Lesson 7) writing scores to a time-series table.
from prometheus_client import Counter, Histogram, Gauge, start_http_server
start_http_server(9000)
LATENCY = Histogram("rag_retrieval_latency_seconds", "Latency",
buckets=[0.01, 0.05, 0.1, 0.25, 0.5, 1.0])
NO_RESULTS = Counter("rag_no_results_total", "Zero-result queries")
TOP_SCORE = Gauge("rag_top_chunk_score", "Top chunk cosine similarity")
Wrap retrieval calls: record latency, increment NO_RESULTS when empty, set TOP_SCORE on hits.
Alert thresholds:
| Metric | Alert threshold | Cause |
|---|---|---|
| Top chunk similarity | < 0.55 (10% of queries) | New topic not yet indexed |
| No-results rate | > 5% | Document gap or corruption |
| Faithfulness | < 0.75 rolling 1-hr avg | Stale index or regression |
| Retrieval P95 latency | > 500 ms | Index growth, HNSW untuned |
Single-node stores become the bottleneck around 1-5 million chunks or hundreds of QPS -- three levers: shard, replicate, batch ingest.
Vector store options at scale:
| Database | Sharding | Managed |
|---|---|---|
| Chroma | No | No |
| pgvector | Via partitioning | Yes (RDS, Supabase) |
| Qdrant | Yes (built-in) | Yes (Qdrant Cloud) |
| Pinecone | Yes (automatic) | Yes |
Batch ingestion -- one-chunk-at-a-time embedding is the most common performance mistake; batch encoding is 10-50x faster.
embed_model = SentenceTransformer("BAAI/bge-small-en-v1.5") # loaded once
def batch_ingest(chunks, metas, coll, client, bs=256):
for i in range(0, len(chunks), bs):
b, m = chunks[i:i+bs], metas[i:i+bs]
vecs = embed_model.encode(b, batch_size=64, normalize_embeddings=True)
pts = [PointStruct(id=str(uuid.uuid4()), vector=vecs[j].tolist(),
payload={"text": b[j], **m[j]}) for j in range(len(b))]
client.upsert(collection_name=coll, points=pts, wait=False)
wait=False returns immediately while Qdrant indexes in the background. For pgvector, use USING hnsw ... WITH (m = 16, ef_construction = 64) and tune ef_search per session.
Production RAG has failure modes absent from prototypes.
| Failure mode | Symptom | Recovery |
|---|---|---|
| Stale data | Answers reference outdated facts | Re-index affected documents |
| Index corruption | Top scores collapse near zero | Restore from snapshot; rebuild |
| Model mismatch | Retrieval degrades after deploy | Roll back; blue/green migration (Step 4) |
| Cache poisoning | Wrong answers at high hit rate | Flush answer cache; investigate |
| Chunk regression | MRR drops on regression suite | Roll back config; re-run Lesson 7 eval |
Graceful degradation -- return a static fallback when retrieval fails or top-chunk score is below threshold:
FALLBACK = "Unable to retrieve relevant documents. Please try again."
def safe_rag_query(query, collection, generate_fn):
try:
hits = instrumented_retrieve(query, collection)
if not hits or hits[0]["score"] < 0.45: # calibrate via Lesson 7 eval
return {"answer": FALLBACK, "status": "no_results"}
return {"answer": answer_with_cache(query, lambda q: generate_fn(q, hits)),
"status": "ok"}
except Exception as exc:
return {"answer": FALLBACK, "status": "error", "detail": str(exc)}
Questions & Answers
Key Takeaways
- Cache at every layer -- embedding, query result, and answer caches each target a different cost; instrument them separately
- Content-hash your documents -- re-index only what changed and handle deletions explicitly so stale chunks cannot outlive their source
- Version the index -- model ID, chunking config, and document revision must be logged so you can replay any query against any historical state
- Monitor for drift, not uptime -- top chunk score, no-results rate, and sampled faithfulness catch quality problems before users do
- Batch ingest and tune HNSW -- one-chunk-at-a-time embedding is the most common performance mistake; fix it before adding hardware
- Test recovery paths -- runbooks are only as good as the last drill; game-day tests turn theoretical plans into practised reflexes
Next Steps: Lesson 10: Multi-Modal RAG