Troubleshooting
Reference
intermediate
Quick Diagnostic Sequence
# 1. Does the answer exist in any chunk? (0 = ingestion problem)
print(sum(1 for c in all_chunks if "keyword" in c.page_content.lower()), "chunks match")
# 2. Are the right chunks retrieved?
for doc, score in vectorstore.similarity_search_with_score(query, k=5):
print(f"{score:.4f} | {doc.metadata.get('source')} | {doc.page_content[:80]}")
# 3. Does the LLM use the context?
chain.verbose = True; chain.invoke({"question": query})
Retrieval Returns Irrelevant Chunks
| Cause | Symptom | Fix |
|---|---|---|
| Query/document vocabulary mismatch | Results topically related but factually wrong | Add hybrid search: dense vectors + BM25 |
| Symmetric model on asymmetric task | Short queries rank poorly against long chunks | Use an instruction-tuned model (multilingual-e5-large-instruct) |
| Wrong domain | Specialist terms return generic results | Use a domain-tuned model; benchmark on MTEB domain subsets |
k too small |
Correct chunk is position 4–8, never seen | Retrieve k=10, re-rank to top-3 with a cross-encoder |
| Metadata filter too tight | Zero results even when docs exist | Log filter hit count; print doc.metadata to verify field names |
from langchain.retrievers import BM25Retriever, EnsembleRetriever
from sentence_transformers import CrossEncoder
ensemble = EnsembleRetriever(
retrievers=[BM25Retriever.from_documents(docs, k=5),
vectorstore.as_retriever(search_kwargs={"k": 5})],
weights=[0.4, 0.6],
)
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")
def rerank(query, candidates, top_n=3):
scores = reranker.predict([(query, d.page_content) for d in candidates])
return [d for _, d in sorted(zip(scores, candidates), reverse=True)[:top_n]]
Answers Ignore Retrieved Context
| Cause | Fix |
|---|---|
| No grounding instruction in system prompt | Add "answer ONLY from context; say 'I don't know' if absent" |
| Context placed after the question | Move context before the question |
| Middle chunks ignored ("lost in the middle") | Place highest-relevance chunk first and last |
| Temperature too high | Set temperature=0 for factual tasks |
| Prompt silently truncated | Count tokens; drop lowest-ranked chunks first |
SYSTEM = "Answer ONLY from context. If absent say 'I don't have enough information.'\nContext: {context}"
import tiktoken
enc = tiktoken.encoding_for_model("gpt-4o")
def trim_to_budget(chunks, system, question, limit=120_000):
used = len(enc.encode(system + question)); out = []
for c in chunks:
cost = len(enc.encode(c.page_content))
if used + cost > limit: break
out.append(c); used += cost
return out
Hallucination Despite Retrieval
| Cause | Fix |
|---|---|
| Chunks are topically close but factually thin | Improve recall (larger k, query expansion); measure faithfulness |
| Old and new document versions both indexed | Delete all chunks for a source before reinserting on update |
| Near-duplicate chunks amplify a wrong topic | Hash chunk content at ingest; skip duplicates |
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy
from datasets import Dataset
# faithfulness < 0.85 → claims not grounded in retrieved context
result = evaluate(
Dataset.from_dict({"question": [q], "answer": [a], "contexts": [[d.page_content for d in docs]]}),
metrics=[faithfulness, answer_relevancy])
import hashlib
def dedup(chunks):
seen, out = set(), []
for c in chunks:
h = hashlib.md5(c.page_content.strip().encode()).hexdigest()
if h not in seen: seen.add(h); out.append(c)
return out
Chunk Boundaries Split Key Facts
| Cause | Fix |
|---|---|
| Fixed-size splitter cuts mid-sentence | Use RecursiveCharacterTextSplitter with 128-token overlap |
| Heading not attached to chunk | Prepend the nearest parent heading to each chunk |
| Tables split across chunks | Extract tables atomically; store each table + caption as one chunk |
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain.retrievers import ParentDocumentRetriever
from langchain.storage import InMemoryStore
# Search small child chunks; return their larger parent passages
retriever = ParentDocumentRetriever(
vectorstore=vectorstore, docstore=InMemoryStore(),
child_splitter=RecursiveCharacterTextSplitter(chunk_size=200),
parent_splitter=RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=128),
)
retriever.add_documents(docs)
Slow Queries at Scale
| Cause | Fix |
|---|---|
| Flat brute-force index | Switch to HNSW — Chroma, pgvector, Qdrant, Pinecone all support it natively |
HNSW ef_search too high |
Lower it; accept a small recall drop for better p95 latency |
| Query embedding on CPU | Move to GPU or use a hosted API with connection pooling |
| No pre-filter — full index scanned | Apply indexed metadata filters (tenant_id, doc_type) before vector scan |
# pgvector — run once after bulk load
# CREATE INDEX ON document_chunks
# USING hnsw (embedding vector_cosine_ops) WITH (m=16, ef_construction=64);
# Qdrant — set HNSW at collection creation
# hnsw_config=HnswConfigDiff(m=16, ef_construct=100)
Stale Indexes After Updates
| Cause | Fix |
|---|---|
| Re-index not triggered after an edit | Track source → sha256; delete-and-reinsert on hash change |
| Old and new chunks coexist | Delete chunks by source before reinserting — most stores do not update in-place |
| Embedding model swapped mid-project | Rebuild the entire index; vectors from different models are not compatible |
from langchain_text_splitters import RecursiveCharacterTextSplitter
import json, hashlib, pathlib
def upsert(vs, doc_id, text, source):
vs._collection.delete(where={"doc_id": doc_id}) # Chroma delete-before-insert
splitter = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=128)
vs.add_documents(splitter.create_documents([text], metadatas=[{"doc_id": doc_id, "source": source}]))
def sync_dir(doc_dir, vs, mf="manifest.json"):
p = pathlib.Path(mf); m = json.loads(p.read_text()) if p.exists() else {}
for path in pathlib.Path(doc_dir).rglob("*.md"):
text = path.read_text(); key = str(path)
h = hashlib.sha256(text.encode()).hexdigest()
if m.get(key) != h: upsert(vs, key, text, key); m[key] = h
p.write_text(json.dumps(m, indent=2))
Store embedding_model in chunk metadata. When you change models, build a new collection in parallel, validate recall, then switch traffic — no in-place upgrade exists.
Evaluation Metrics at a Glance
| Metric | What it measures | Target |
|---|---|---|
| Precision@k | Fraction of top-k chunks that are relevant | > 0.6 at k=5 |
| Recall@k | Fraction of all relevant chunks in top-k | > 0.7 at k=10 |
| MRR | Reciprocal rank of first relevant result | > 0.7 |
| Faithfulness (RAGAS) | Answer claims supported by context | > 0.85 |
| Answer Relevancy (RAGAS) | Answer addresses the question asked | > 0.80 |
| Context Recall (RAGAS) | Ground-truth facts covered by retrieved chunks | > 0.75 |