Troubleshooting

Reference intermediate

Quick Diagnostic Sequence

# 1. Does the answer exist in any chunk? (0 = ingestion problem)
print(sum(1 for c in all_chunks if "keyword" in c.page_content.lower()), "chunks match")
# 2. Are the right chunks retrieved?
for doc, score in vectorstore.similarity_search_with_score(query, k=5):
    print(f"{score:.4f} | {doc.metadata.get('source')} | {doc.page_content[:80]}")
# 3. Does the LLM use the context?
chain.verbose = True; chain.invoke({"question": query})

Retrieval Returns Irrelevant Chunks

Cause Symptom Fix
Query/document vocabulary mismatch Results topically related but factually wrong Add hybrid search: dense vectors + BM25
Symmetric model on asymmetric task Short queries rank poorly against long chunks Use an instruction-tuned model (multilingual-e5-large-instruct)
Wrong domain Specialist terms return generic results Use a domain-tuned model; benchmark on MTEB domain subsets
k too small Correct chunk is position 4–8, never seen Retrieve k=10, re-rank to top-3 with a cross-encoder
Metadata filter too tight Zero results even when docs exist Log filter hit count; print doc.metadata to verify field names
from langchain.retrievers import BM25Retriever, EnsembleRetriever
from sentence_transformers import CrossEncoder

ensemble = EnsembleRetriever(
    retrievers=[BM25Retriever.from_documents(docs, k=5),
                vectorstore.as_retriever(search_kwargs={"k": 5})],
    weights=[0.4, 0.6],
)

reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")
def rerank(query, candidates, top_n=3):
    scores = reranker.predict([(query, d.page_content) for d in candidates])
    return [d for _, d in sorted(zip(scores, candidates), reverse=True)[:top_n]]

Answers Ignore Retrieved Context

Cause Fix
No grounding instruction in system prompt Add "answer ONLY from context; say 'I don't know' if absent"
Context placed after the question Move context before the question
Middle chunks ignored ("lost in the middle") Place highest-relevance chunk first and last
Temperature too high Set temperature=0 for factual tasks
Prompt silently truncated Count tokens; drop lowest-ranked chunks first
SYSTEM = "Answer ONLY from context. If absent say 'I don't have enough information.'\nContext: {context}"

import tiktoken
enc = tiktoken.encoding_for_model("gpt-4o")
def trim_to_budget(chunks, system, question, limit=120_000):
    used = len(enc.encode(system + question)); out = []
    for c in chunks:
        cost = len(enc.encode(c.page_content))
        if used + cost > limit: break
        out.append(c); used += cost
    return out

Hallucination Despite Retrieval

Cause Fix
Chunks are topically close but factually thin Improve recall (larger k, query expansion); measure faithfulness
Old and new document versions both indexed Delete all chunks for a source before reinserting on update
Near-duplicate chunks amplify a wrong topic Hash chunk content at ingest; skip duplicates
from ragas import evaluate
from ragas.metrics import faithfulness, answer_relevancy
from datasets import Dataset
# faithfulness < 0.85 → claims not grounded in retrieved context
result = evaluate(
    Dataset.from_dict({"question": [q], "answer": [a], "contexts": [[d.page_content for d in docs]]}),
    metrics=[faithfulness, answer_relevancy])

import hashlib
def dedup(chunks):
    seen, out = set(), []
    for c in chunks:
        h = hashlib.md5(c.page_content.strip().encode()).hexdigest()
        if h not in seen: seen.add(h); out.append(c)
    return out

Chunk Boundaries Split Key Facts

Cause Fix
Fixed-size splitter cuts mid-sentence Use RecursiveCharacterTextSplitter with 128-token overlap
Heading not attached to chunk Prepend the nearest parent heading to each chunk
Tables split across chunks Extract tables atomically; store each table + caption as one chunk
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain.retrievers import ParentDocumentRetriever
from langchain.storage import InMemoryStore

# Search small child chunks; return their larger parent passages
retriever = ParentDocumentRetriever(
    vectorstore=vectorstore, docstore=InMemoryStore(),
    child_splitter=RecursiveCharacterTextSplitter(chunk_size=200),
    parent_splitter=RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=128),
)
retriever.add_documents(docs)

Slow Queries at Scale

Cause Fix
Flat brute-force index Switch to HNSW — Chroma, pgvector, Qdrant, Pinecone all support it natively
HNSW ef_search too high Lower it; accept a small recall drop for better p95 latency
Query embedding on CPU Move to GPU or use a hosted API with connection pooling
No pre-filter — full index scanned Apply indexed metadata filters (tenant_id, doc_type) before vector scan
# pgvector — run once after bulk load
# CREATE INDEX ON document_chunks
# USING hnsw (embedding vector_cosine_ops) WITH (m=16, ef_construction=64);

# Qdrant — set HNSW at collection creation
# hnsw_config=HnswConfigDiff(m=16, ef_construct=100)

Stale Indexes After Updates

Cause Fix
Re-index not triggered after an edit Track source → sha256; delete-and-reinsert on hash change
Old and new chunks coexist Delete chunks by source before reinserting — most stores do not update in-place
Embedding model swapped mid-project Rebuild the entire index; vectors from different models are not compatible
from langchain_text_splitters import RecursiveCharacterTextSplitter
import json, hashlib, pathlib

def upsert(vs, doc_id, text, source):
    vs._collection.delete(where={"doc_id": doc_id})   # Chroma delete-before-insert
    splitter = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=128)
    vs.add_documents(splitter.create_documents([text], metadatas=[{"doc_id": doc_id, "source": source}]))

def sync_dir(doc_dir, vs, mf="manifest.json"):
    p = pathlib.Path(mf); m = json.loads(p.read_text()) if p.exists() else {}
    for path in pathlib.Path(doc_dir).rglob("*.md"):
        text = path.read_text(); key = str(path)
        h = hashlib.sha256(text.encode()).hexdigest()
        if m.get(key) != h: upsert(vs, key, text, key); m[key] = h
    p.write_text(json.dumps(m, indent=2))

Store embedding_model in chunk metadata. When you change models, build a new collection in parallel, validate recall, then switch traffic — no in-place upgrade exists.


Evaluation Metrics at a Glance

Metric What it measures Target
Precision@k Fraction of top-k chunks that are relevant > 0.6 at k=5
Recall@k Fraction of all relevant chunks in top-k > 0.7 at k=10
MRR Reciprocal rank of first relevant result > 0.7
Faithfulness (RAGAS) Answer claims supported by context > 0.85
Answer Relevancy (RAGAS) Answer addresses the question asked > 0.80
Context Recall (RAGAS) Ground-truth facts covered by retrieved chunks > 0.75

Related References