Evaluation & Quality
Learning Outcomes
- Calculate and interpret precision@k, recall@k, and MRR to measure retrieval component quality
- Measure generation quality using faithfulness and answer relevance with the RAGAS framework
- Build a golden test set that covers your domain's edge cases and failure modes
- Implement an LLM-as-judge scorer for scalable automated evaluation
- Design a continuous evaluation loop that catches quality regressions before they reach production
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why you cannot improve what you cannot measure |
| Explain | 7 min | Retrieval metrics: precision@k, recall@k, MRR |
| Demo | 8 min | Computing retrieval metrics in Python |
| Explain | 7 min | Generation metrics: faithfulness, relevance, correctness |
| Demo | 10 min | Running RAGAS on your pipeline |
| Explain | 7 min | Building a golden test set and LLM-as-judge |
| Demo | 5 min | Wiring evaluation into a CI gate |
| Wrap-up | 3 min | Choosing the right metric for the right question |
Before You Begin
Pre-work:
- Complete Lesson 5: Building a Basic RAG Pipeline — you need a working pipeline to evaluate
- Complete Lesson 6: Advanced Retrieval — retrieval metrics are most useful after you have techniques to compare
- Optionally review Lesson 2: Embeddings Explained for background on similarity measures used inside metrics
Shopping List:
- Python 3.10+ environment with your RAG pipeline from Lessons 5–6
pip install ragas langchain-openai datasets(see Step 4 for alternatives)- An LLM API key (OpenAI, Anthropic, or any LangChain-compatible provider) — used by the LLM-as-judge evaluator
- 10–20 representative questions about your corpus to seed the golden set (prose notes are fine to start)
Before writing a single line of eval code, it helps to understand what makes RAG evaluation genuinely difficult. There are three independent components that can each fail silently:
| Component | What it does | How it fails |
|---|---|---|
| Retriever | Fetches relevant chunks from the vector store | Returns plausible-sounding but wrong documents |
| Generator | Synthesises an answer from retrieved context | Ignores context, hallucinates, or over-hedges |
| Pipeline integration | Assembles prompt, truncates context, routes queries | Good retrieval + good generation still produces bad answers if the prompt is malformed |
A retriever that scores 0.90 precision@5 can still doom your pipeline if the generator ignores the top-ranked chunk. Conversely, a highly capable generator will hallucinate confidently when the retriever gives it nothing useful. You need metrics at each stage to know which component to fix.
The discipline is called offline evaluation: you run your pipeline against a fixed test set of question–answer pairs, score the results, and track those scores over time. Done well, it gives you a regression gate you can attach to CI, so a bad embedding model swap or chunking change trips an alert before it reaches users.
Retrieval evaluation requires a relevance judgement for each query: for query Q, which documents in your corpus are actually relevant? With those labels, you can compute the metrics below. In practice you obtain labels either from human annotation or, for bootstrapping, from an LLM that reads each chunk and the query.
Precision@k
Precision@k answers: of the k chunks I retrieved, what fraction were actually relevant?
precision@k = (relevant chunks in top-k) / k
A pipeline with k=5 that returns 3 relevant chunks has precision@5 = 0.6. It does not reward finding all relevant chunks — only that the ones you surface are good. Use precision@k when your generator has a tight context window and you can only afford to pass a few chunks.
Recall@k
Recall@k answers: of all the relevant chunks that exist, what fraction did I retrieve in the top k?
recall@k = (relevant chunks in top-k) / (total relevant chunks in corpus)
If your corpus has 10 relevant chunks for a query and you retrieve 4 of them in your top 5, recall@5 = 0.4. Use recall@k when your query requires synthesising across multiple source chunks — missing any one degrades the answer.
Mean Reciprocal Rank (MRR)
MRR is rank-aware. For each query it computes the reciprocal of the rank of the first relevant result, then averages across queries:
MRR = (1 / |Q|) * sum(1 / rank_i for each query i)
If the first relevant chunk appears at rank 1 it contributes 1.0; at rank 2, 0.5; at rank 3, 0.33. MRR drops sharply when relevance is buried. It is the right metric when you pass only the top-1 chunk to your generator, or when your UX shows a single best answer.
Which metric to choose?
| Situation | Recommended metric |
|---|---|
| Tight context, top-1 or top-2 passed to LLM | MRR, Precision@1 |
| Multi-chunk synthesis (summarisation, comparison) | Recall@k |
| Balanced retrieval quality signal | Precision@5 + Recall@5 together |
| Ranking quality matters (re-ranker tuning) | NDCG@k (rank-weighted recall) |
The code below is self-contained and assumes you have a retrieve(query, k) function that returns a list of chunk dicts with an id field. Replace it with your actual retriever.
from dataclasses import dataclass, field
from typing import List, Set
@dataclass
class RetrievalSample:
query: str
relevant_ids: Set[str] # ground-truth relevant chunk IDs
retrieved_ids: List[str] # ordered list returned by your retriever
def precision_at_k(sample: RetrievalSample, k: int) -> float:
top_k = set(sample.retrieved_ids[:k])
hits = top_k & sample.relevant_ids
return len(hits) / k
def recall_at_k(sample: RetrievalSample, k: int) -> float:
if not sample.relevant_ids:
return 0.0
top_k = set(sample.retrieved_ids[:k])
hits = top_k & sample.relevant_ids
return len(hits) / len(sample.relevant_ids)
def reciprocal_rank(sample: RetrievalSample) -> float:
for rank, chunk_id in enumerate(sample.retrieved_ids, start=1):
if chunk_id in sample.relevant_ids:
return 1.0 / rank
return 0.0
def evaluate_retrieval(samples: List[RetrievalSample], k: int = 5) -> dict:
p_scores = [precision_at_k(s, k) for s in samples]
r_scores = [recall_at_k(s, k) for s in samples]
rr_scores = [reciprocal_rank(s) for s in samples]
return {
f"precision@{k}": round(sum(p_scores) / len(p_scores), 4),
f"recall@{k}": round(sum(r_scores) / len(r_scores), 4),
"MRR": round(sum(rr_scores) / len(rr_scores), 4),
}
# --- Example usage ---
# Assume your retriever is wired up; build samples from your golden set
golden = [
RetrievalSample(
query="What is the refund policy for digital products?",
relevant_ids={"chunk_042", "chunk_043"},
retrieved_ids=["chunk_043", "chunk_001", "chunk_042", "chunk_099", "chunk_017"],
),
RetrievalSample(
query="How do I upgrade my subscription tier?",
relevant_ids={"chunk_077"},
retrieved_ids=["chunk_077", "chunk_022", "chunk_031", "chunk_009", "chunk_055"],
),
]
metrics = evaluate_retrieval(golden, k=5)
print(metrics)
# {"precision@5": 0.5, "recall@5": 1.0, "MRR": 0.75}
The key insight: you need chunk IDs to exist in both your golden set and your vector store. When you build your golden set (Step 5), record the chunk ID alongside the question so you can compute these scores.
Once the retriever has fetched good chunks, you need to verify the generator actually uses them. The three metrics below cover complementary failure modes.
Faithfulness
Faithfulness measures whether every factual claim in the generated answer is supported by the retrieved context. It does not require the answer to be complete — only that nothing it asserts contradicts or goes beyond the context.
faithfulness = (claims in the answer supported by context) / (total claims in the answer)
A faithfulness score of 1.0 means the model grounded every claim. A score of 0.5 means half the claims are hallucinated. Low faithfulness signals that your prompt is not constraining the model tightly enough, or the retrieved context is too thin for the question.
Answer Relevance (Response Relevancy)
Answer relevance measures whether the answer actually addresses the user's question. A model that responds with accurate but off-topic information scores high on faithfulness but low on relevance. RAGAS computes this by having an LLM generate hypothetical questions from the response, then measuring the cosine similarity between those hypothetical questions and the original query.
Factual Correctness
Factual correctness compares the generated answer against a reference answer, measuring overlap of factual claims. This is an end-to-end metric: it catches cases where the retriever surfaced correct documents but the generator misstated the key fact.
Metric failure modes at a glance
| Symptom | Likely cause | Metric to watch |
|---|---|---|
| Model invents details not in context | Generator hallucination | Faithfulness |
| Answer wanders off-topic | Prompt lacks instruction to stay on task | Answer relevance |
| Answer is grounded but factually wrong | Retriever returning wrong-but-similar docs | Factual correctness |
| Answer is correct but incomplete | Low recall@k | Recall@k + factual correctness together |
RAGAS (Retrieval-Augmented Generation Assessment) is an open-source evaluation framework that automates generation-side metrics using an LLM-as-judge internally.
Install it alongside a LangChain-compatible LLM wrapper:
pip install ragas langchain-openai datasets
The core data structure is an EvaluationDataset built from SingleTurnSample objects. Each sample carries the user's question, the retrieved contexts your pipeline returned, the generated response, and (optionally) a reference answer for correctness scoring.
import os
from ragas import EvaluationDataset, evaluate
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import LLMContextRecall, Faithfulness, FactualCorrectness
from langchain_openai import ChatOpenAI
# Wrap your LLM for use as the judge
evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini"))
# Build your evaluation dataset
# Each row is one query your pipeline answered, with the retrieved chunks
# and the generated response recorded verbatim.
raw_samples = [
{
"user_input": "What is the return policy for software licenses?",
"retrieved_contexts": [
"Software licenses are non-refundable once the activation key has been used.",
"For unused keys purchased within 14 days, a full refund is available.",
],
"response": "Software licenses cannot be refunded after the key is used. "
"Unused keys may be refunded within 14 days of purchase.",
"reference": "Software licenses are non-refundable once activated; "
"unused keys qualify for a refund within 14 days.",
},
{
"user_input": "Does the Pro plan include API access?",
"retrieved_contexts": [
"The Pro plan includes unlimited API calls at up to 1000 req/min.",
],
"response": "Yes, the Pro plan includes API access with a rate limit of 1000 requests per minute.",
"reference": "The Pro plan includes API access at 1000 requests per minute.",
},
]
dataset = EvaluationDataset.from_list(raw_samples)
# Run evaluation — each metric uses the judge LLM internally
result = evaluate(
dataset=dataset,
metrics=[
LLMContextRecall(), # retrieval: did context cover the reference answer?
Faithfulness(), # generation: are all claims grounded in context?
FactualCorrectness(), # end-to-end: does the answer match the reference?
],
llm=evaluator_llm,
)
print(result)
# {'llm_context_recall': 0.97, 'faithfulness': 1.0, 'factual_correctness': 0.88}
# Export to a DataFrame for tracking over time
df = result.to_pandas()
df.to_csv("eval_results.csv", index=False)
The to_pandas() call is important: save results to CSV or a database on every evaluation run so you can track trends. A single score means nothing; a score that dropped 0.08 between last week and today tells you exactly where to look.
A golden test set is a curated collection of question–answer pairs that represent the real queries your system will face. It is the most important investment you can make in evaluation quality — a bad golden set produces meaningless metrics. The practical minimum for a useful signal is 50 questions; 100–200 gives stable averages.
What makes a good golden question?
Cover these categories deliberately:
| Category | Example | Why it matters |
|---|---|---|
| Factual lookup | "What is the SLA for Priority 1 tickets?" | Tests basic retrieval precision |
| Multi-chunk synthesis | "Compare the refund policies for hardware vs. software" | Tests recall + coherence |
| Edge case / negation | "Is there a free trial for the Enterprise plan?" (answer: no) | Tests that model doesn't hallucinate |
| Ambiguous query | "How do I cancel?" (cancel what?) | Tests graceful handling of underspecified input |
| Out-of-scope query | "What is the capital of France?" | Tests that model admits ignorance |
Collecting questions with a generation assist
You can bootstrap questions from your corpus, then have humans verify and correct them:
import json
from openai import OpenAI
client = OpenAI()
def generate_questions_from_chunk(chunk_text: str, n: int = 3) -> list[dict]:
"""
Given a chunk, ask an LLM to generate plausible questions
a user might ask that this chunk answers.
Returns a list of dicts with 'question' and 'reference_answer'.
"""
prompt = f"""Given the following document excerpt, write {n} realistic questions
a user might ask whose answer is found in this excerpt.
For each question, also write a concise reference answer (1-2 sentences).
Excerpt:
{chunk_text}
Return a JSON array of objects with keys "question" and "reference_answer".
Return only the JSON array, no surrounding text."""
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"},
)
# Parse and return; wrap in try/except in production
data = json.loads(response.choices[0].message.content)
# Model returns a JSON object with a key containing the array
for key, value in data.items():
if isinstance(value, list):
return value
return []
# Usage: sample chunks from your corpus and generate candidates
sample_chunk = (
"The Enterprise plan includes dedicated infrastructure, a 99.99% uptime SLA, "
"and a designated Customer Success Manager. There is no free trial; "
"contact sales for a proof-of-concept engagement."
)
candidates = generate_questions_from_chunk(sample_chunk, n=3)
print(json.dumps(candidates, indent=2))
# [
# {"question": "Does the Enterprise plan offer a free trial?",
# "reference_answer": "No, the Enterprise plan does not include a free trial."},
# ...
# ]
After generation, have a human review each candidate: delete questions that are trivially answered outside your corpus, correct reference answers where the model was imprecise, and add edge-case questions the generator missed. Store the final set as JSONL:
{"question": "Does the Enterprise plan offer a free trial?", "reference_answer": "No, there is no free trial for the Enterprise plan. Contact sales for a proof-of-concept engagement.", "relevant_chunk_ids": ["chunk_017"]}
{"question": "What uptime SLA does the Enterprise plan guarantee?", "reference_answer": "The Enterprise plan carries a 99.99% uptime SLA.", "relevant_chunk_ids": ["chunk_017"]}
RAGAS covers general-purpose metrics well. For domain-specific scoring — "does this answer follow our support tone guidelines?" or "does it correctly cite the relevant policy section?" — you need a custom judge.
The pattern is straightforward: write a scoring prompt, call an LLM, parse the score. Keeping the output format strict (JSON with a numeric score and a brief reason) makes parsing reliable and gives you an audit trail.
import json
import re
from openai import OpenAI
from dataclasses import dataclass
client = OpenAI()
@dataclass
class JudgeResult:
score: float # 0.0 to 1.0
reason: str
FAITHFULNESS_PROMPT = """You are an expert evaluator assessing whether an AI-generated answer
is faithful to the provided context — meaning every factual claim in the answer is
explicitly supported by the context and nothing is invented.
Score the answer on a scale of 0 to 1:
1.0 = every claim is directly supported by the context
0.5 = some claims are supported, others are not
0.0 = the answer contradicts or ignores the context entirely
Return a JSON object with keys "score" (float, 0-1) and "reason" (one sentence).
Context:
{context}
Question:
{question}
Answer:
{answer}"""
def judge_faithfulness(context: str, question: str, answer: str) -> JudgeResult:
prompt = FAITHFULNESS_PROMPT.format(
context=context, question=question, answer=answer
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"},
temperature=0, # deterministic scoring
)
raw = json.loads(response.choices[0].message.content)
return JudgeResult(score=float(raw["score"]), reason=raw["reason"])
# Usage
result = judge_faithfulness(
context="The Basic plan costs $9/month and supports up to 3 users.",
question="How much does the Basic plan cost?",
answer="The Basic plan is $9 per month and supports up to 3 users.",
)
print(result)
# JudgeResult(score=1.0, reason='Both claims are directly supported by the context.')
Calibrating your judge
Before trusting the judge's scores in CI, calibrate it against a small human-annotated set:
# Calibration set: 20-30 (context, question, answer, human_score) tuples
calibration = [
{
"context": "Refunds are processed within 5 business days.",
"question": "How long do refunds take?",
"answer": "Refunds take 5 business days.",
"human_score": 1.0,
},
{
"context": "Refunds are processed within 5 business days.",
"question": "How long do refunds take?",
"answer": "Refunds typically take 2-3 business days.", # hallucinated
"human_score": 0.0,
},
]
errors = []
for item in calibration:
result = judge_faithfulness(
item["context"], item["question"], item["answer"]
)
errors.append(abs(result.score - item["human_score"]))
mean_abs_error = sum(errors) / len(errors)
print(f"Judge MAE vs human: {mean_abs_error:.3f}")
# Target: MAE < 0.15 before using in CI
temperature=0 makes LLM judges more consistent but not perfectly reproducible across model versions. Re-run calibration whenever the judge model is updated — small model changes can shift scores by 0.05–0.10.Evaluation only pays off if it runs automatically. The pattern below stores results in a SQLite database (easy to swap for Postgres in production), computes a pass/fail gate, and exits non-zero on failure — making it a natural CI step.
import sqlite3
import datetime
import sys
import json
# Assume `run_pipeline_on_golden_set` returns a list of completed sample dicts
# with all fields needed for RAGAS + your retrieval metrics.
DB_PATH = "eval_results.db"
THRESHOLDS = {
"faithfulness": 0.85,
"llm_context_recall": 0.80,
"factual_correctness":0.75,
"precision_at_5": 0.70,
"mrr": 0.80,
}
def init_db(path: str):
conn = sqlite3.connect(path)
conn.execute("""
CREATE TABLE IF NOT EXISTS runs (
id INTEGER PRIMARY KEY AUTOINCREMENT,
run_at TEXT,
git_sha TEXT,
metrics TEXT, -- JSON blob
passed INTEGER
)
""")
conn.commit()
return conn
def record_run(conn, git_sha: str, metrics: dict, passed: bool):
conn.execute(
"INSERT INTO runs (run_at, git_sha, metrics, passed) VALUES (?, ?, ?, ?)",
(
datetime.datetime.utcnow().isoformat(),
git_sha,
json.dumps(metrics),
int(passed),
),
)
conn.commit()
def evaluate_and_gate(metrics: dict) -> bool:
failures = []
for key, threshold in THRESHOLDS.items():
value = metrics.get(key)
if value is None:
print(f" SKIP {key}: not present in results")
continue
status = "PASS" if value >= threshold else "FAIL"
symbol = "+" if status == "PASS" else "x"
print(f" [{symbol}] {key}: {value:.3f} (threshold {threshold:.2f})")
if status == "FAIL":
failures.append(key)
return len(failures) == 0
if __name__ == "__main__":
import subprocess
git_sha = subprocess.check_output(
["git", "rev-parse", "--short", "HEAD"], text=True
).strip()
# Replace with your actual pipeline evaluation
# from your_pipeline import run_pipeline_on_golden_set
# ragas_results, retrieval_results = run_pipeline_on_golden_set("golden.jsonl")
# metrics = {**ragas_results, **retrieval_results}
# Stub for illustration
metrics = {
"faithfulness": 0.92,
"llm_context_recall": 0.84,
"factual_correctness": 0.78,
"precision_at_5": 0.73,
"mrr": 0.81,
}
print(f"Evaluation run for {git_sha}")
passed = evaluate_and_gate(metrics)
conn = init_db(DB_PATH)
record_run(conn, git_sha, metrics, passed)
if not passed:
print("\nEvaluation gate FAILED. See failures above.")
sys.exit(1)
print("\nEvaluation gate PASSED.")
sys.exit(0)
Add this as a CI step in your GitHub Actions or equivalent workflow. In the YAML you pass the API key from your repository secrets (the exact secrets syntax is described in the GitHub Actions docs — use secrets.OPENAI_API_KEY as the env value):
# .github/workflows/rag-eval.yml (fragment)
- name: Run RAG evaluation gate
run: python eval/run_eval.py
env:
OPENAI_API_KEY: "<your-secret-ref-here>"
The SQLite history lets you plot metric trends with a simple query:
SELECT run_at, json_extract(metrics, '$.faithfulness') AS faithfulness,
json_extract(metrics, '$.mrr') AS mrr
FROM runs
ORDER BY run_at DESC
LIMIT 20;
Questions & Answers
evaluate() function already parallelises internally. Fourth, cache judge outputs keyed on (context, question, answer) hash — if the pipeline produces the same answer for the same input, there is no need to re-judge. On 200 samples with GPT-4o-mini at current rates, a full RAGAS run costs well under a dollar.Key Takeaways
- Evaluate each stage independently. Precision@k, recall@k, and MRR measure the retriever; faithfulness and answer relevance measure the generator; factual correctness measures the end-to-end pipeline. A problem in any one stage can hide behind strong scores in the others.
- The golden set is the foundation. Fifty to two hundred human-verified question–answer pairs that cover your domain's edge cases are more valuable than any particular metric. Build it once and keep it under version control.
- RAGAS automates generation-side evaluation using LLM-as-judge for faithfulness, context recall, and factual correctness. It integrates with LangChain and LlamaIndex and outputs a DataFrame you can track over time.
- LLM judges are useful but biased. Calibrate your judge against human-scored examples and use relative score trends for regression gates rather than trusting absolute numbers as ground truth.
- Wire evaluation into CI. An evaluation script that exits non-zero on threshold failure catches regressions before deployment. Start with soft thresholds below your current baseline, observe variance, then tighten.
- High faithfulness + low correctness = retrieval problem. The most common misdiagnosis in RAG evaluation is blaming the generator for a retriever failure. Always check retrieval metrics first when end-to-end quality drops.
Next Steps: Lesson 8: Knowledge Graphs + RAG