Production RAG

55 min advanced Lesson 9

Learning Outcomes

  • Design a multi-layer caching strategy that cuts retrieval latency and LLM spend
  • Implement incremental re-indexing so document updates propagate without a full rebuild
  • Version your index and trace which document revision produced any given answer
  • Build monitoring that surfaces retrieval-quality drift before users notice it
  • Diagnose and recover from the failure modes unique to production RAG

Lesson Plan

Segment Duration Topic
Intro 3 min What changes when you go to production
Steps 1-2 14 min Caching at every layer
Step 3 9 min Document updates and re-indexing
Step 4 7 min Versioning the index
Step 5 9 min Monitoring for quality drift
Step 6 8 min Scaling and batch ingestion
Step 7 2 min Failure modes quick reference
Wrap-up 3 min Production operations plan

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ with pip install chromadb qdrant-client sentence-transformers redis hiredis prometheus-client psycopg2-binary
  • Redis: docker run -p 6379:6379 redis:7-alpine
  • An LLM API key (Anthropic, OpenAI, or local)

1 The Production Gap

A prototype RAG pipeline has one job: retrieve documents and generate an answer. Production adds five more.

Prototype assumption Production reality
Queries are unique 60-80% repeat within an hour
Documents are static Docs change daily; deletions without notice
One model forever Model changes force migrations
Errors are obvious Quality drifts slowly and silently
Low QPS, single node Hundreds of concurrent requests

Three layers to build on Lesson 5: Cache (embedding, query results, answers), Index (versioned, model-tagged), Ops (monitoring, alerting, rollback).

NOTE
Build ops before launch
Most teams hit production problems after a silent week of drift. Building observability before launch is cheaper than debugging under user pressure.

2 Caching at Every Layer

RAG has three expensive operations -- embedding (~10-50 ms), vector search (~20-200 ms), LLM generation (~500 ms+) -- each cacheable with separate TTLs.

Embedding cache (1 hr) -- SHA-256 the query, store raw float32 bytes in Redis. Hit: np.frombuffer(r.get(key), dtype=np.float32). Miss: embed, then r.setex(key, 3600, vec.tobytes()).

Query result cache (5 min) -- key on the vector rounded to 3 decimal places; store the serialised hit list as JSON.

Answer cache (30 min) -- only for stable queries; skip for time-relative or user-specific ones.

Cache layer Key Hit rate TTL
Embedding SHA-256 of query text 70-90% 1 hr
Query results Rounded vector hash 40-60% 5 min
Answer Normalised query string 20-50% 30 min
WARNING
Invalidate on re-index
When you re-index a document, flush cache entries that referenced its chunks. Keep a reverse index from chunk_id to cache keys and delete on update. If too complex, use a short TTL instead.

3 Document Updates and Re-indexing

Nightly full rebuilds break down at scale. The production pattern: re-embed only documents whose content changed.

Content-hash gating -- SHA-256 the text; if the hash matches the stored record, skip re-embedding entirely. When 5% of docs change daily, this cuts ingestion cost by 95%.

import hashlib

def update_document(doc_id, text, collection, registry, chunker, embedder):
    h = hashlib.sha256(text.encode()).hexdigest()
    if doc_id in registry and registry[doc_id]["hash"] == h:
        return  # unchanged
    if doc_id in registry:
        collection.delete(ids=registry[doc_id]["chunk_ids"])
    chunks = chunker(text)
    ids = [f"{doc_id}::{i}" for i in range(len(chunks))]
    collection.add(ids=ids, documents=chunks,
                   embeddings=[embedder(c).tolist() for c in chunks],
                   metadatas=[{"doc_id": doc_id, "i": i} for i in range(len(chunks))])
    registry[doc_id] = {"hash": h, "chunk_ids": ids,
                        "version": registry.get(doc_id, {}).get("version", 0) + 1}

Deletion handling -- skipping leaves deleted docs in the index forever. Call collection.delete(ids=registry[doc_id]["chunk_ids"]) then del registry[doc_id].

NOTE
Soft deletes for compliance
Add a deleted_at timestamp to chunk metadata and filter at query time instead of hard-deleting -- proves what was available at any past date.

4 Versioning the Index

Versioning answers the question that has no answer without it: "Last Tuesday's answer was wrong -- what documents did my system retrieve?"

What to version:

Artifact How
Document registry Append-only SQL event log
Embedding model ID Model ID in every chunk's metadata
Chunking config chunk_size and overlap in chunk metadata

Minimal PostgreSQL event log:

CREATE TABLE rag_events (
    event_id BIGSERIAL PRIMARY KEY, occurred_at TIMESTAMPTZ DEFAULT NOW(),
    event_type TEXT NOT NULL, doc_id TEXT, doc_version INT,
    content_hash TEXT, model_id TEXT, chunk_count INT
);
CREATE INDEX ON rag_events (doc_id, occurred_at DESC);

One row per update_document call lets you answer: at time T, what version of doc X was indexed and which model produced its vectors?

Embedding model migration -- new model vectors are geometrically incompatible with old ones. Safe procedure: build a parallel collection, shadow-test at 5%, compare precision@5 and MRR, ramp to 100%, retire the old.

WARNING
Never mix embedding models in one collection
Embedding document A with model X and document B with model Y means B is invisible to queries using model X -- its vectors live in a different geometric space. Always store and validate the model ID before searching.

5 Monitoring for Quality Drift

Quality drift kills user trust slowly -- new topics appear in questions before documents are indexed, or popular docs are deleted without replacement.

Expose three Prometheus metrics (scrape with Grafana) and add a 1% faithfulness-sampling loop (LLM-as-judge, Lesson 7) writing scores to a time-series table.

from prometheus_client import Counter, Histogram, Gauge, start_http_server
start_http_server(9000)
LATENCY = Histogram("rag_retrieval_latency_seconds", "Latency",
                    buckets=[0.01, 0.05, 0.1, 0.25, 0.5, 1.0])
NO_RESULTS = Counter("rag_no_results_total", "Zero-result queries")
TOP_SCORE  = Gauge("rag_top_chunk_score", "Top chunk cosine similarity")

Wrap retrieval calls: record latency, increment NO_RESULTS when empty, set TOP_SCORE on hits.

Alert thresholds:

Metric Alert threshold Cause
Top chunk similarity < 0.55 (10% of queries) New topic not yet indexed
No-results rate > 5% Document gap or corruption
Faithfulness < 0.75 rolling 1-hr avg Stale index or regression
Retrieval P95 latency > 500 ms Index growth, HNSW untuned
TIP
Store concrete examples
Aggregate metrics say something is wrong; stored (query, chunk IDs, answer) samples show you why. Also track no-results queries separately -- they identify exactly which user intents your corpus does not cover.

6 Scaling: Sharding, Replication, and Batch Ingestion

Single-node stores become the bottleneck around 1-5 million chunks or hundreds of QPS -- three levers: shard, replicate, batch ingest.

Vector store options at scale:

Database Sharding Managed
Chroma No No
pgvector Via partitioning Yes (RDS, Supabase)
Qdrant Yes (built-in) Yes (Qdrant Cloud)
Pinecone Yes (automatic) Yes

Batch ingestion -- one-chunk-at-a-time embedding is the most common performance mistake; batch encoding is 10-50x faster.

embed_model = SentenceTransformer("BAAI/bge-small-en-v1.5")  # loaded once

def batch_ingest(chunks, metas, coll, client, bs=256):
    for i in range(0, len(chunks), bs):
        b, m = chunks[i:i+bs], metas[i:i+bs]
        vecs = embed_model.encode(b, batch_size=64, normalize_embeddings=True)
        pts = [PointStruct(id=str(uuid.uuid4()), vector=vecs[j].tolist(),
                           payload={"text": b[j], **m[j]}) for j in range(len(b))]
        client.upsert(collection_name=coll, points=pts, wait=False)

wait=False returns immediately while Qdrant indexes in the background. For pgvector, use USING hnsw ... WITH (m = 16, ef_construction = 64) and tune ef_search per session.

NOTE
Shard vs replicate
Shard when the index exceeds RAM (~5M 384-dim vectors at ~7 GB). Replicate for throughput and availability. Start with replication; add sharding only when measured.

7 Failure Modes and Recovery

Production RAG has failure modes absent from prototypes.

Failure mode Symptom Recovery
Stale data Answers reference outdated facts Re-index affected documents
Index corruption Top scores collapse near zero Restore from snapshot; rebuild
Model mismatch Retrieval degrades after deploy Roll back; blue/green migration (Step 4)
Cache poisoning Wrong answers at high hit rate Flush answer cache; investigate
Chunk regression MRR drops on regression suite Roll back config; re-run Lesson 7 eval

Graceful degradation -- return a static fallback when retrieval fails or top-chunk score is below threshold:

FALLBACK = "Unable to retrieve relevant documents. Please try again."

def safe_rag_query(query, collection, generate_fn):
    try:
        hits = instrumented_retrieve(query, collection)
        if not hits or hits[0]["score"] < 0.45:  # calibrate via Lesson 7 eval
            return {"answer": FALLBACK, "status": "no_results"}
        return {"answer": answer_with_cache(query, lambda q: generate_fn(q, hits)),
                "status": "ok"}
    except Exception as exc:
        return {"answer": FALLBACK, "status": "error", "detail": str(exc)}
TIP
Write a runbook before launch
Map each alert to a response: who gets paged, what commands to run, what rollback looks like. Writing it forces you to design recovery before you need it.
WARNING
Test recovery paths
Run a game-day drill: corrupt the index, flush the cache, swap the embedding model. A procedure never exercised will fail when needed.

Questions & Answers

Q: Our corpus updates hundreds of times per minute. Can we stay in sync without constant re-embedding?
Yes -- the content-hash check (Step 3) means you only re-embed changed documents; batch encoding (Step 6) handles those efficiently. For very high rates, use a hot tier (recent docs, always current) plus a cold tier (historical, batch-updated).
Q: Our embedding model vendor is deprecating the model we use. How do we migrate safely?
Build a parallel collection with the replacement model, shadow-test at 5% traffic, validate precision@5 and MRR, then cut over gradually. Never serve mixed-model results from one collection -- vectors from different models are geometrically incompatible.
Q: Users say answers were correct last week but wrong now. We have not deployed anything. Where do we look?
In order: (1) source documents changed -- query your version log; (2) cache poisoning -- flush the answer cache and re-test; (3) upstream returned different content briefly -- check your registry for indexed-at timestamp gaps; (4) embedding API silently updated weights -- compare a test query vector from before and after the window.
Q: How do we pick cache TTLs when we do not know how fast documents change?
Instrument your registry to track time-between-updates. If 90% of documents are unchanged in 24 hours, a 1-hour answer TTL works. If documents update every 5 minutes, skip the answer cache. Calibrate by domain -- compliance systems often tolerate zero staleness.
Q: Latency is fine at baseline but spikes under load. HNSW is tuned. What else?
Three likely causes: (1) CPU-bound embedding -- offload to a dedicated service and use the embedding cache (Step 2); (2) ef_search too high for concurrency -- lower it and compensate with re-ranking (Lesson 6); (3) connection pool exhaustion -- share the vector store client, don't instantiate per-request.

Key Takeaways

  1. Cache at every layer -- embedding, query result, and answer caches each target a different cost; instrument them separately
  2. Content-hash your documents -- re-index only what changed and handle deletions explicitly so stale chunks cannot outlive their source
  3. Version the index -- model ID, chunking config, and document revision must be logged so you can replay any query against any historical state
  4. Monitor for drift, not uptime -- top chunk score, no-results rate, and sampled faithfulness catch quality problems before users do
  5. Batch ingest and tune HNSW -- one-chunk-at-a-time embedding is the most common performance mistake; fix it before adding hardware
  6. Test recovery paths -- runbooks are only as good as the last drill; game-day tests turn theoretical plans into practised reflexes

Next Steps: Lesson 10: Multi-Modal RAG