Evaluation & Quality

50 min advanced Lesson 7

Learning Outcomes

  • Calculate and interpret precision@k, recall@k, and MRR to measure retrieval component quality
  • Measure generation quality using faithfulness and answer relevance with the RAGAS framework
  • Build a golden test set that covers your domain's edge cases and failure modes
  • Implement an LLM-as-judge scorer for scalable automated evaluation
  • Design a continuous evaluation loop that catches quality regressions before they reach production

Lesson Plan

Segment Duration Topic
Intro 3 min Why you cannot improve what you cannot measure
Explain 7 min Retrieval metrics: precision@k, recall@k, MRR
Demo 8 min Computing retrieval metrics in Python
Explain 7 min Generation metrics: faithfulness, relevance, correctness
Demo 10 min Running RAGAS on your pipeline
Explain 7 min Building a golden test set and LLM-as-judge
Demo 5 min Wiring evaluation into a CI gate
Wrap-up 3 min Choosing the right metric for the right question

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ environment with your RAG pipeline from Lessons 5–6
  • pip install ragas langchain-openai datasets (see Step 4 for alternatives)
  • An LLM API key (OpenAI, Anthropic, or any LangChain-compatible provider) — used by the LLM-as-judge evaluator
  • 10–20 representative questions about your corpus to seed the golden set (prose notes are fine to start)

1 Why Evaluation Is Hard — and Why You Need It

Before writing a single line of eval code, it helps to understand what makes RAG evaluation genuinely difficult. There are three independent components that can each fail silently:

Component What it does How it fails
Retriever Fetches relevant chunks from the vector store Returns plausible-sounding but wrong documents
Generator Synthesises an answer from retrieved context Ignores context, hallucinates, or over-hedges
Pipeline integration Assembles prompt, truncates context, routes queries Good retrieval + good generation still produces bad answers if the prompt is malformed

A retriever that scores 0.90 precision@5 can still doom your pipeline if the generator ignores the top-ranked chunk. Conversely, a highly capable generator will hallucinate confidently when the retriever gives it nothing useful. You need metrics at each stage to know which component to fix.

The discipline is called offline evaluation: you run your pipeline against a fixed test set of question–answer pairs, score the results, and track those scores over time. Done well, it gives you a regression gate you can attach to CI, so a bad embedding model swap or chunking change trips an alert before it reaches users.

NOTE
Evaluation vs. Monitoring
Offline evaluation (this lesson) uses a golden test set you control. Production monitoring (covered in Lesson 9) streams live traffic. Both are necessary: offline catches regressions before deploy, monitoring catches distribution shift after deploy.

2 Retrieval Metrics: Precision@k, Recall@k, and MRR

Retrieval evaluation requires a relevance judgement for each query: for query Q, which documents in your corpus are actually relevant? With those labels, you can compute the metrics below. In practice you obtain labels either from human annotation or, for bootstrapping, from an LLM that reads each chunk and the query.

Precision@k

Precision@k answers: of the k chunks I retrieved, what fraction were actually relevant?

precision@k = (relevant chunks in top-k) / k

A pipeline with k=5 that returns 3 relevant chunks has precision@5 = 0.6. It does not reward finding all relevant chunks — only that the ones you surface are good. Use precision@k when your generator has a tight context window and you can only afford to pass a few chunks.

Recall@k

Recall@k answers: of all the relevant chunks that exist, what fraction did I retrieve in the top k?

recall@k = (relevant chunks in top-k) / (total relevant chunks in corpus)

If your corpus has 10 relevant chunks for a query and you retrieve 4 of them in your top 5, recall@5 = 0.4. Use recall@k when your query requires synthesising across multiple source chunks — missing any one degrades the answer.

Mean Reciprocal Rank (MRR)

MRR is rank-aware. For each query it computes the reciprocal of the rank of the first relevant result, then averages across queries:

MRR = (1 / |Q|) * sum(1 / rank_i  for each query i)

If the first relevant chunk appears at rank 1 it contributes 1.0; at rank 2, 0.5; at rank 3, 0.33. MRR drops sharply when relevance is buried. It is the right metric when you pass only the top-1 chunk to your generator, or when your UX shows a single best answer.

Which metric to choose?

Situation Recommended metric
Tight context, top-1 or top-2 passed to LLM MRR, Precision@1
Multi-chunk synthesis (summarisation, comparison) Recall@k
Balanced retrieval quality signal Precision@5 + Recall@5 together
Ranking quality matters (re-ranker tuning) NDCG@k (rank-weighted recall)
TIP
Quick Sanity Check
Run precision@1 first. If it is below 0.7 on your golden set, your embedding model or chunking strategy is the bottleneck — no amount of re-ranking or prompt engineering will fix a retriever that is wrong most of the time.

3 Computing Retrieval Metrics in Python

The code below is self-contained and assumes you have a retrieve(query, k) function that returns a list of chunk dicts with an id field. Replace it with your actual retriever.

from dataclasses import dataclass, field
from typing import List, Set


@dataclass
class RetrievalSample:
    query: str
    relevant_ids: Set[str]      # ground-truth relevant chunk IDs
    retrieved_ids: List[str]    # ordered list returned by your retriever


def precision_at_k(sample: RetrievalSample, k: int) -> float:
    top_k = set(sample.retrieved_ids[:k])
    hits = top_k & sample.relevant_ids
    return len(hits) / k


def recall_at_k(sample: RetrievalSample, k: int) -> float:
    if not sample.relevant_ids:
        return 0.0
    top_k = set(sample.retrieved_ids[:k])
    hits = top_k & sample.relevant_ids
    return len(hits) / len(sample.relevant_ids)


def reciprocal_rank(sample: RetrievalSample) -> float:
    for rank, chunk_id in enumerate(sample.retrieved_ids, start=1):
        if chunk_id in sample.relevant_ids:
            return 1.0 / rank
    return 0.0


def evaluate_retrieval(samples: List[RetrievalSample], k: int = 5) -> dict:
    p_scores = [precision_at_k(s, k) for s in samples]
    r_scores = [recall_at_k(s, k) for s in samples]
    rr_scores = [reciprocal_rank(s) for s in samples]
    return {
        f"precision@{k}": round(sum(p_scores) / len(p_scores), 4),
        f"recall@{k}":    round(sum(r_scores) / len(r_scores), 4),
        "MRR":            round(sum(rr_scores) / len(rr_scores), 4),
    }


# --- Example usage ---
# Assume your retriever is wired up; build samples from your golden set
golden = [
    RetrievalSample(
        query="What is the refund policy for digital products?",
        relevant_ids={"chunk_042", "chunk_043"},
        retrieved_ids=["chunk_043", "chunk_001", "chunk_042", "chunk_099", "chunk_017"],
    ),
    RetrievalSample(
        query="How do I upgrade my subscription tier?",
        relevant_ids={"chunk_077"},
        retrieved_ids=["chunk_077", "chunk_022", "chunk_031", "chunk_009", "chunk_055"],
    ),
]

metrics = evaluate_retrieval(golden, k=5)
print(metrics)
# {"precision@5": 0.5, "recall@5": 1.0, "MRR": 0.75}

The key insight: you need chunk IDs to exist in both your golden set and your vector store. When you build your golden set (Step 5), record the chunk ID alongside the question so you can compute these scores.

WARNING
Chunk ID Stability
If you re-chunk your corpus, chunk IDs change and your golden set becomes invalid. Store golden sets against stable document IDs + character offsets, then re-derive the chunk IDs after each re-index. Alternatively, use fuzzy text matching to re-align.

4 Generation Metrics: Faithfulness, Answer Relevance, and Factual Correctness

Once the retriever has fetched good chunks, you need to verify the generator actually uses them. The three metrics below cover complementary failure modes.

Faithfulness

Faithfulness measures whether every factual claim in the generated answer is supported by the retrieved context. It does not require the answer to be complete — only that nothing it asserts contradicts or goes beyond the context.

faithfulness = (claims in the answer supported by context) / (total claims in the answer)

A faithfulness score of 1.0 means the model grounded every claim. A score of 0.5 means half the claims are hallucinated. Low faithfulness signals that your prompt is not constraining the model tightly enough, or the retrieved context is too thin for the question.

Answer Relevance (Response Relevancy)

Answer relevance measures whether the answer actually addresses the user's question. A model that responds with accurate but off-topic information scores high on faithfulness but low on relevance. RAGAS computes this by having an LLM generate hypothetical questions from the response, then measuring the cosine similarity between those hypothetical questions and the original query.

Factual Correctness

Factual correctness compares the generated answer against a reference answer, measuring overlap of factual claims. This is an end-to-end metric: it catches cases where the retriever surfaced correct documents but the generator misstated the key fact.

Metric failure modes at a glance

Symptom Likely cause Metric to watch
Model invents details not in context Generator hallucination Faithfulness
Answer wanders off-topic Prompt lacks instruction to stay on task Answer relevance
Answer is grounded but factually wrong Retriever returning wrong-but-similar docs Factual correctness
Answer is correct but incomplete Low recall@k Recall@k + factual correctness together
NOTE
Faithfulness vs. Correctness
Faithfulness is relative to the retrieved context. Correctness is relative to ground truth. A system can be faithful (it only asserts what the context says) but incorrect (the context was about the wrong thing). Always measure both.

5 Running RAGAS on Your Pipeline

RAGAS (Retrieval-Augmented Generation Assessment) is an open-source evaluation framework that automates generation-side metrics using an LLM-as-judge internally.

Install it alongside a LangChain-compatible LLM wrapper:

pip install ragas langchain-openai datasets

The core data structure is an EvaluationDataset built from SingleTurnSample objects. Each sample carries the user's question, the retrieved contexts your pipeline returned, the generated response, and (optionally) a reference answer for correctness scoring.

import os
from ragas import EvaluationDataset, evaluate
from ragas.llms import LangchainLLMWrapper
from ragas.metrics import LLMContextRecall, Faithfulness, FactualCorrectness
from langchain_openai import ChatOpenAI

# Wrap your LLM for use as the judge
evaluator_llm = LangchainLLMWrapper(ChatOpenAI(model="gpt-4o-mini"))

# Build your evaluation dataset
# Each row is one query your pipeline answered, with the retrieved chunks
# and the generated response recorded verbatim.
raw_samples = [
    {
        "user_input":        "What is the return policy for software licenses?",
        "retrieved_contexts": [
            "Software licenses are non-refundable once the activation key has been used.",
            "For unused keys purchased within 14 days, a full refund is available.",
        ],
        "response":  "Software licenses cannot be refunded after the key is used. "
                     "Unused keys may be refunded within 14 days of purchase.",
        "reference": "Software licenses are non-refundable once activated; "
                     "unused keys qualify for a refund within 14 days.",
    },
    {
        "user_input":        "Does the Pro plan include API access?",
        "retrieved_contexts": [
            "The Pro plan includes unlimited API calls at up to 1000 req/min.",
        ],
        "response":  "Yes, the Pro plan includes API access with a rate limit of 1000 requests per minute.",
        "reference": "The Pro plan includes API access at 1000 requests per minute.",
    },
]

dataset = EvaluationDataset.from_list(raw_samples)

# Run evaluation — each metric uses the judge LLM internally
result = evaluate(
    dataset=dataset,
    metrics=[
        LLMContextRecall(),    # retrieval: did context cover the reference answer?
        Faithfulness(),        # generation: are all claims grounded in context?
        FactualCorrectness(),  # end-to-end: does the answer match the reference?
    ],
    llm=evaluator_llm,
)

print(result)
# {'llm_context_recall': 0.97, 'faithfulness': 1.0, 'factual_correctness': 0.88}

# Export to a DataFrame for tracking over time
df = result.to_pandas()
df.to_csv("eval_results.csv", index=False)

The to_pandas() call is important: save results to CSV or a database on every evaluation run so you can track trends. A single score means nothing; a score that dropped 0.08 between last week and today tells you exactly where to look.

TIP
Cheap Judge Models Work
GPT-4o-mini and equivalent smaller models score faithfulness and relevance with accuracy close to GPT-4o at a fraction of the cost. Reserve the larger model for difficult factual correctness judgements where nuance matters.
WARNING
LLM Judge Bias
LLM-as-judge exhibits systematic biases: it prefers longer responses, shows positional bias (favouring the first claim listed), and tends to score its own outputs higher. Calibrate your judge against a small set of human-verified examples before trusting its absolute scores — use relative trends for regression testing, not raw scores as ground truth.

6 Building a Golden Test Set

A golden test set is a curated collection of question–answer pairs that represent the real queries your system will face. It is the most important investment you can make in evaluation quality — a bad golden set produces meaningless metrics. The practical minimum for a useful signal is 50 questions; 100–200 gives stable averages.

What makes a good golden question?

Cover these categories deliberately:

Category Example Why it matters
Factual lookup "What is the SLA for Priority 1 tickets?" Tests basic retrieval precision
Multi-chunk synthesis "Compare the refund policies for hardware vs. software" Tests recall + coherence
Edge case / negation "Is there a free trial for the Enterprise plan?" (answer: no) Tests that model doesn't hallucinate
Ambiguous query "How do I cancel?" (cancel what?) Tests graceful handling of underspecified input
Out-of-scope query "What is the capital of France?" Tests that model admits ignorance

Collecting questions with a generation assist

You can bootstrap questions from your corpus, then have humans verify and correct them:

import json
from openai import OpenAI

client = OpenAI()

def generate_questions_from_chunk(chunk_text: str, n: int = 3) -> list[dict]:
    """
    Given a chunk, ask an LLM to generate plausible questions
    a user might ask that this chunk answers.
    Returns a list of dicts with 'question' and 'reference_answer'.
    """
    prompt = f"""Given the following document excerpt, write {n} realistic questions
a user might ask whose answer is found in this excerpt.
For each question, also write a concise reference answer (1-2 sentences).

Excerpt:
{chunk_text}

Return a JSON array of objects with keys "question" and "reference_answer".
Return only the JSON array, no surrounding text."""

    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        response_format={"type": "json_object"},
    )
    # Parse and return; wrap in try/except in production
    data = json.loads(response.choices[0].message.content)
    # Model returns a JSON object with a key containing the array
    for key, value in data.items():
        if isinstance(value, list):
            return value
    return []


# Usage: sample chunks from your corpus and generate candidates
sample_chunk = (
    "The Enterprise plan includes dedicated infrastructure, a 99.99% uptime SLA, "
    "and a designated Customer Success Manager. There is no free trial; "
    "contact sales for a proof-of-concept engagement."
)

candidates = generate_questions_from_chunk(sample_chunk, n=3)
print(json.dumps(candidates, indent=2))
# [
#   {"question": "Does the Enterprise plan offer a free trial?",
#    "reference_answer": "No, the Enterprise plan does not include a free trial."},
#   ...
# ]

After generation, have a human review each candidate: delete questions that are trivially answered outside your corpus, correct reference answers where the model was imprecise, and add edge-case questions the generator missed. Store the final set as JSONL:

{"question": "Does the Enterprise plan offer a free trial?", "reference_answer": "No, there is no free trial for the Enterprise plan. Contact sales for a proof-of-concept engagement.", "relevant_chunk_ids": ["chunk_017"]}
{"question": "What uptime SLA does the Enterprise plan guarantee?", "reference_answer": "The Enterprise plan carries a 99.99% uptime SLA.", "relevant_chunk_ids": ["chunk_017"]}
TIP
Versioning the Golden Set
Commit your golden set to the same repository as your pipeline code. When you update it, record why — new edge cases, corpus expansion, user-reported failures. The history is as valuable as the set itself.

7 Custom LLM-as-Judge for Domain-Specific Scoring

RAGAS covers general-purpose metrics well. For domain-specific scoring — "does this answer follow our support tone guidelines?" or "does it correctly cite the relevant policy section?" — you need a custom judge.

The pattern is straightforward: write a scoring prompt, call an LLM, parse the score. Keeping the output format strict (JSON with a numeric score and a brief reason) makes parsing reliable and gives you an audit trail.

import json
import re
from openai import OpenAI
from dataclasses import dataclass

client = OpenAI()


@dataclass
class JudgeResult:
    score: float      # 0.0 to 1.0
    reason: str


FAITHFULNESS_PROMPT = """You are an expert evaluator assessing whether an AI-generated answer
is faithful to the provided context — meaning every factual claim in the answer is
explicitly supported by the context and nothing is invented.

Score the answer on a scale of 0 to 1:
  1.0 = every claim is directly supported by the context
  0.5 = some claims are supported, others are not
  0.0 = the answer contradicts or ignores the context entirely

Return a JSON object with keys "score" (float, 0-1) and "reason" (one sentence).

Context:
{context}

Question:
{question}

Answer:
{answer}"""


def judge_faithfulness(context: str, question: str, answer: str) -> JudgeResult:
    prompt = FAITHFULNESS_PROMPT.format(
        context=context, question=question, answer=answer
    )
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": prompt}],
        response_format={"type": "json_object"},
        temperature=0,  # deterministic scoring
    )
    raw = json.loads(response.choices[0].message.content)
    return JudgeResult(score=float(raw["score"]), reason=raw["reason"])


# Usage
result = judge_faithfulness(
    context="The Basic plan costs $9/month and supports up to 3 users.",
    question="How much does the Basic plan cost?",
    answer="The Basic plan is $9 per month and supports up to 3 users.",
)
print(result)
# JudgeResult(score=1.0, reason='Both claims are directly supported by the context.')

Calibrating your judge

Before trusting the judge's scores in CI, calibrate it against a small human-annotated set:

# Calibration set: 20-30 (context, question, answer, human_score) tuples
calibration = [
    {
        "context":      "Refunds are processed within 5 business days.",
        "question":     "How long do refunds take?",
        "answer":       "Refunds take 5 business days.",
        "human_score":  1.0,
    },
    {
        "context":      "Refunds are processed within 5 business days.",
        "question":     "How long do refunds take?",
        "answer":       "Refunds typically take 2-3 business days.",  # hallucinated
        "human_score":  0.0,
    },
]

errors = []
for item in calibration:
    result = judge_faithfulness(
        item["context"], item["question"], item["answer"]
    )
    errors.append(abs(result.score - item["human_score"]))

mean_abs_error = sum(errors) / len(errors)
print(f"Judge MAE vs human: {mean_abs_error:.3f}")
# Target: MAE < 0.15 before using in CI
WARNING
Temperature Zero Is Not Deterministic
Setting temperature=0 makes LLM judges more consistent but not perfectly reproducible across model versions. Re-run calibration whenever the judge model is updated — small model changes can shift scores by 0.05–0.10.

8 Wiring Evaluation into a Continuous Quality Gate

Evaluation only pays off if it runs automatically. The pattern below stores results in a SQLite database (easy to swap for Postgres in production), computes a pass/fail gate, and exits non-zero on failure — making it a natural CI step.

import sqlite3
import datetime
import sys
import json

# Assume `run_pipeline_on_golden_set` returns a list of completed sample dicts
# with all fields needed for RAGAS + your retrieval metrics.

DB_PATH = "eval_results.db"
THRESHOLDS = {
    "faithfulness":       0.85,
    "llm_context_recall": 0.80,
    "factual_correctness":0.75,
    "precision_at_5":     0.70,
    "mrr":                0.80,
}


def init_db(path: str):
    conn = sqlite3.connect(path)
    conn.execute("""
        CREATE TABLE IF NOT EXISTS runs (
            id         INTEGER PRIMARY KEY AUTOINCREMENT,
            run_at     TEXT,
            git_sha    TEXT,
            metrics    TEXT,   -- JSON blob
            passed     INTEGER
        )
    """)
    conn.commit()
    return conn


def record_run(conn, git_sha: str, metrics: dict, passed: bool):
    conn.execute(
        "INSERT INTO runs (run_at, git_sha, metrics, passed) VALUES (?, ?, ?, ?)",
        (
            datetime.datetime.utcnow().isoformat(),
            git_sha,
            json.dumps(metrics),
            int(passed),
        ),
    )
    conn.commit()


def evaluate_and_gate(metrics: dict) -> bool:
    failures = []
    for key, threshold in THRESHOLDS.items():
        value = metrics.get(key)
        if value is None:
            print(f"  SKIP  {key}: not present in results")
            continue
        status = "PASS" if value >= threshold else "FAIL"
        symbol = "+" if status == "PASS" else "x"
        print(f"  [{symbol}] {key}: {value:.3f} (threshold {threshold:.2f})")
        if status == "FAIL":
            failures.append(key)
    return len(failures) == 0


if __name__ == "__main__":
    import subprocess

    git_sha = subprocess.check_output(
        ["git", "rev-parse", "--short", "HEAD"], text=True
    ).strip()

    # Replace with your actual pipeline evaluation
    # from your_pipeline import run_pipeline_on_golden_set
    # ragas_results, retrieval_results = run_pipeline_on_golden_set("golden.jsonl")
    # metrics = {**ragas_results, **retrieval_results}

    # Stub for illustration
    metrics = {
        "faithfulness":        0.92,
        "llm_context_recall":  0.84,
        "factual_correctness": 0.78,
        "precision_at_5":      0.73,
        "mrr":                 0.81,
    }

    print(f"Evaluation run for {git_sha}")
    passed = evaluate_and_gate(metrics)

    conn = init_db(DB_PATH)
    record_run(conn, git_sha, metrics, passed)

    if not passed:
        print("\nEvaluation gate FAILED. See failures above.")
        sys.exit(1)

    print("\nEvaluation gate PASSED.")
    sys.exit(0)

Add this as a CI step in your GitHub Actions or equivalent workflow. In the YAML you pass the API key from your repository secrets (the exact secrets syntax is described in the GitHub Actions docs — use secrets.OPENAI_API_KEY as the env value):

# .github/workflows/rag-eval.yml  (fragment)
- name: Run RAG evaluation gate
  run: python eval/run_eval.py
  env:
    OPENAI_API_KEY: "<your-secret-ref-here>"

The SQLite history lets you plot metric trends with a simple query:

SELECT run_at, json_extract(metrics, '$.faithfulness') AS faithfulness,
               json_extract(metrics, '$.mrr')          AS mrr
FROM runs
ORDER BY run_at DESC
LIMIT 20;
TIP
Start with Soft Gates
In the first week, set thresholds 10–15 points below your current scores so the gate never fires. Run it for a week to build intuition for natural score variance, then tighten thresholds to just below your observed baseline. Hard gates on unstable metrics waste engineering time.
NOTE
Framework Alternatives
RAGAS is the most widely-adopted open-source option. DeepEval offers a unit-test-style API that integrates with pytest — useful if you prefer assertion-based tests over metric averaging. TruLens has a strong dashboard for experiment comparison. All three are active in 2026 and complement each other: RAGAS for metric exploration, DeepEval for CI gates, TruLens for dashboards.

Questions & Answers

Q: Our golden set only has 30 questions. Are the metric averages actually meaningful?
Thirty questions gives you very wide confidence intervals — a single bad retrieval swing can move precision@5 by ±0.03. That said, 30 samples is enough to detect large regressions (a new embedding model that drops precision by 0.10) and catches obvious bugs. For stable production gates, aim for 100+ questions. Prioritise diversity over volume: 30 carefully chosen questions covering distinct intents are more useful than 100 variations of the same lookup query. Track standard deviation alongside the mean — if std > 0.15, your corpus may be too heterogeneous for a single threshold.
Q: RAGAS calls the judge LLM for every sample. On 200 samples that gets expensive. How do we keep costs down?
Several strategies help. First, use a smaller judge model (GPT-4o-mini or equivalent) for faithfulness and context recall — these metrics do not require deep reasoning, and smaller models perform comparably. Second, run the full suite only on pull requests that touch the pipeline; use a 20-sample smoke-test subset on every commit. Third, batch your RAGAS calls: the evaluate() function already parallelises internally. Fourth, cache judge outputs keyed on (context, question, answer) hash — if the pipeline produces the same answer for the same input, there is no need to re-judge. On 200 samples with GPT-4o-mini at current rates, a full RAGAS run costs well under a dollar.
Q: Faithfulness is high but users still complain the answers are wrong. What is going on?
Faithfulness only measures whether claims are grounded in retrieved context. If the retriever is returning the wrong chunks — plausible but outdated or off-topic documents — a faithful answer to bad context is still a wrong answer. Check factual correctness against your golden reference answers, and check retrieval precision@5 for the failing queries. The pattern "high faithfulness, low correctness" is almost always a retrieval problem. Run your failing queries through the pipeline manually, log the retrieved chunks, and see whether the correct source document was retrieved at all. If it was not in the top-5, that is a retrieval gap; if it was retrieved but not used, that is a prompt or context-assembly problem.
Q: We re-chunked our corpus and now all our golden chunk IDs are invalid. Do we have to rebuild the whole golden set?
Not from scratch. Store golden sets against stable identifiers — document ID plus character offset range — rather than chunk IDs. After re-chunking, run a script that finds which new chunks overlap each old offset range (a simple text-matching pass works) and re-derives the relevant chunk IDs. If you did not capture offsets originally, the fastest recovery is to take your golden questions, run them through the new retriever, and have a human (or LLM judge) verify which of the top-5 returned chunks are actually relevant. That rebuilds the relevance judgements in an hour or two without having to re-write the questions or reference answers.
Q: Should we evaluate on our training data or hold out a separate test set?
If your golden set was used to tune any component — chunk size, embedding model choice, re-ranker threshold — it is contaminated as a test set and will overestimate quality. Keep a small held-out test set (20–30 questions) that is never used for tuning decisions, and do not look at its scores during active development. Use the larger development golden set for iteration, and reserve the test set for final milestone validation or major architecture comparisons. The discipline matters: the moment you start optimising for the test set, it stops measuring generalisation.

Key Takeaways

  1. Evaluate each stage independently. Precision@k, recall@k, and MRR measure the retriever; faithfulness and answer relevance measure the generator; factual correctness measures the end-to-end pipeline. A problem in any one stage can hide behind strong scores in the others.
  2. The golden set is the foundation. Fifty to two hundred human-verified question–answer pairs that cover your domain's edge cases are more valuable than any particular metric. Build it once and keep it under version control.
  3. RAGAS automates generation-side evaluation using LLM-as-judge for faithfulness, context recall, and factual correctness. It integrates with LangChain and LlamaIndex and outputs a DataFrame you can track over time.
  4. LLM judges are useful but biased. Calibrate your judge against human-scored examples and use relative score trends for regression gates rather than trusting absolute numbers as ground truth.
  5. Wire evaluation into CI. An evaluation script that exits non-zero on threshold failure catches regressions before deployment. Start with soft thresholds below your current baseline, observe variance, then tighten.
  6. High faithfulness + low correctness = retrieval problem. The most common misdiagnosis in RAG evaluation is blaming the generator for a retriever failure. Always check retrieval metrics first when end-to-end quality drops.

Next Steps: Lesson 8: Knowledge Graphs + RAG