Evaluating Agent Performance

45 min advanced Lesson 8

Learning Outcomes

  • Define a multi-dimensional metric set: completion, efficiency, cost, and reliability
  • Distinguish outcome-based grading from trajectory-based grading and apply each
  • Build a reusable evaluation harness that runs an agent over a labelled task set
  • Interpret the benchmarks SWE-bench, GAIA, and WebArena and their limits
  • Measure run-to-run reliability and detect when an agent fails to recognise it is stuck

Lesson Plan

Segment Duration Topic
Intro 3 min Why accuracy alone lies about agents
Explain 7 min Outcome vs trajectory grading
Build 9 min A minimal eval harness: cost and efficiency
Build 8 min Reliability across repeated runs
Explain 8 min Benchmarks: SWE-bench, GAIA, WebArena
Build 8 min Detecting "stuck" and graceful failure
Wrap-up 2 min Key takeaways and next steps

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ and the anthropic SDK (pip install anthropic), with ANTHROPIC_API_KEY exported
  • The Lesson 6 agent importable as a function (goal in; result + trajectory + usage out)
  • 10-20 tasks you already know the correct answer to

1 Why a Single Accuracy Number Lies

For a classifier, accuracy tells the whole story. For an agent it hides almost everything that matters in production. Two agents can both "solve" 70% of tasks while behaving completely differently:

Agent A Agent B
Task completion 70% 70%
Avg cost per task $0.04 $0.61
Run-to-run consistency 95% 48%
Recognises when stuck Yes No (loops to cap)

Agent A is shippable. Agent B burns fifteen times the budget, can't repeat a success, and grinds against its cap instead of giving up. Completion told you none of that. A useful evaluation tracks at least four orthogonal dimensions — task completion, efficiency (steps vs the optimal path), cost per completed task, and reliability (same input, same outcome?) — plus a fifth, graceful failure: does the agent know when it has failed? And always hold out a test set the agent never influenced, or you measure memorisation, not capability.

NOTE
Key Insight
Agents are stochastic and multi-step, so a single run is a sample, not a measurement. Evaluate distributions across repeated runs, not one lucky trajectory.

2 Outcome Grading vs Trajectory Grading

There are two ways to score a run, and you usually need both. Outcome grading checks the final result against a known answer. A verifier is a function from (task, final_output) to a boolean or score:

def grade_outcome(task: dict, final_output: str) -> bool:
    # Exact-match verifier; real ones run tests or compare fields.
    return final_output.strip().lower() == task["expected_answer"].strip().lower()

For coding agents the verifier runs the repository's unit tests against the produced patch and returns pass/fail. That is exactly how SWE-bench grades, and why it is trusted: no judgement call.

Trajectory grading inspects how the agent got there. Walk the tool_use steps and compute signals: step count, repeated calls (len(names) - len(set(names))), whether a forbidden tool like delete_file appears. Use it when the outcome looks right but the path was unsafe, or when there is no single correct answer. For open-ended tasks, LLM-as-judge has a separate model grade the output against a rubric — it scales but is fallible, so calibrate it against human labels first.

TIP
Title
Prefer a programmatic verifier (tests, schema checks, numeric tolerance) whenever you can build one. Reserve LLM-as-judge for subjective outputs, and spot-check its scores against human grades.

3 A Minimal Harness with Cost and Efficiency

A harness is just a loop: load tasks, run the agent, grade, aggregate. Store tasks as a flat JSON list — each with an id, goal, expected_answer, and optimal_steps (your estimate of the shortest path, the denominator for efficiency). Assume the Lesson 6 agent exposes run_agent(goal) -> AgentResult, returning final text plus trajectory and token usage.

from dataclasses import dataclass

@dataclass
class TaskResult:
    completed: bool; steps: int; optimal_steps: int
    in_tok: int; out_tok: int; error: str | None = None

def run_eval(tasks: list[dict]) -> list[TaskResult]:
    out = []
    for t in tasks:
        try:
            r = run_agent(t["goal"])                 # your Lesson 6 agent
            out.append(TaskResult(
                grade_outcome(t, r.final_output), len(r.trajectory),
                t["optimal_steps"], r.usage.input_tokens, r.usage.output_tokens))
        except Exception as e:                        # a crash is a failed task
            out.append(TaskResult(False, 0, t["optimal_steps"], 0, 0, str(e)))
    return out

The aggregates fall out directly: completion_rate, step_efficiency (mean of optimal_steps / steps, 1.0 is perfect), crash_rate, and the key one, cost_per_success — total spend (including failures) divided by successes, so an agent that fails expensively is correctly punished. Token counts come from the Anthropic API usage block times your model's current per-million-token rates. Track cost_per_success over time, not just completion_rate.

WARNING
Watch Out
An exception inside the agent is a failure, not a data point to drop. Silently skipping crashes makes a brittle agent look more reliable than one that merely answers wrong.

4 Reliability: Run It Many Times

Because agents sample tokens, the same task can pass on Monday and fail on Tuesday. Reliability measures that variance with pass@k and the stricter always-pass rate: run each task k times and look at the distribution.

from collections import defaultdict

def reliability_eval(tasks: list[dict], k: int = 5) -> dict:
    outcomes = defaultdict(list)          # task_id -> [bool, ...]
    for t in tasks:
        for _ in range(k):
            r = run_agent(t["goal"])
            outcomes[t["id"]].append(grade_outcome(t, r.final_output))
    n = len(outcomes)
    return {
        "pass_at_1": sum(o[0] for o in outcomes.values()) / n,
        "pass_at_k": sum(any(o) for o in outcomes.values()) / n,  # any of k
        "always_pass": sum(all(o) for o in outcomes.values()) / n,
    }

Reading the spread diagnoses the problem: high pass@k but low pass@1 means it can solve the task unreliably (tighten prompts, lower temperature, add a verification step); all three roughly equal means it is stable (wrong-but-stable is a capability gap, not variance); low always_pass with high pass@k is the dangerous case — great in a demo, weak under one-shot load.

WARNING
Watch Out
Temperature 0 reduces but does not eliminate variance — tool results, timeouts, and ordering still introduce nondeterminism. Report reliability from real repeated rollouts, and treat pass@1 as the honest production number: users do not get k retries.

5 Standard Benchmarks and How to Read Them

Know the public benchmarks before building everything yourself. Each targets a different skill and grades programmatically — no vibes.

  • SWE-bench evaluates models on real-world GitHub issues: given a codebase and an issue, generate a patch, scored by running the repo's unit tests in an isolated container. SWE-bench Verified is a 500-instance subset engineers confirmed solvable.
  • GAIA is 466 questions needing reasoning, multimodality, browsing, and tool use, across three difficulty levels by steps/tools needed (Level 1 few steps, up to Level 3 long sequences). Headline gap: humans ~92% vs ~15% for GPT-4 with plugins at publication.
  • WebArena is a realistic, self-hostable web environment with functional sites (e-commerce, forums, GitLab, CMS), scoring agents on functional task correctness.
WARNING
Watch Out
Use benchmarks only to choose a base model and sanity-check your stack; trust your own held-out suite for go/no-go. Contamination is real — popular tasks leak into training data and inflate scores — so treat a headline percentage as marketing.

6 Detecting "Stuck": Measuring Graceful Failure

The worst agent behaviour is not failing — it is failing while burning your whole budget and never admitting it. A good agent recognises a dead end and stops. Measure this from the trajectory by classifying each failed run:

def diagnose_failure(result, step_cap: int) -> str:
    traj = result.trajectory
    last = [s["name"] for s in traj if s["type"] == "tool_use"][-4:]
    if len(traj) >= step_cap:
        return "hit_cap"           # never stopped — worst case
    if len(last) == 4 and len(set(last)) == 1:
        return "looping"           # same tool over and over
    if result.gave_up_explicitly:
        return "graceful_giveup"   # best failure: it knew
    return "wrong_but_terminated"  # finished cleanly, just wrong

Report the distribution of these buckets. Failures dominated by graceful_giveup are far safer to deploy than ones stuck at hit_cap or looping, even at identical completion rates. This ties back to the resource limits from Lesson 7 — the step cap is your backstop, but a healthy agent rarely needs it. To push failures into that bucket, give the agent an explicit report_stuck tool to call (with a reason) when blocked.

NOTE
Key Insight
A self-aware failure is a feature. An agent that calls report_stuck after two honest attempts saves money and lets a human step in early — far better than one that thrashes to the cap and returns a confident wrong answer.

Questions & Answers

Q: I don't have ground-truth answers for my real tasks. How do I grade them?
Build a programmatic verifier even without a single answer (valid JSON, required fields, value within tolerance, tests pass). Failing that, use LLM-as-judge against a rubric, calibrated on 30-50 hand-graded examples first. Never ship an eval whose grader you have not audited against human labels.
Q: How many tasks does my eval suite actually need?
Enough that a one-task change does not swing the headline wildly. Twenty gives 5% resolution — fine for catching regressions; for a release decision, want 100+ spanning your real difficulty distribution, stratified so an easy-heavy suite does not hide hard-task failures.
Q: Running every task five times is slow and expensive. Is there a cheaper way?
Split the suite: run the full set once per change for completion and cost, and reserve k-repeat reliability runs for a smaller subset or a nightly job. Use the Anthropic batch API for non-interactive runs, and parallelise across tasks since they are independent.
Q: My agent scores well on a public benchmark but poorly on my product. Why?
Distribution shift (your tasks and formats differ from the benchmark's) and contamination (benchmark tasks may be in training data, inflating the score). A private, held-out, domain-specific suite should drive decisions; use public benchmarks only for coarse model selection.
Q: Should I count partial progress, or is it strictly pass/fail?
Start strict — binary completion is unambiguous and hard to game. Add partial credit only for genuinely separable sub-goals (e.g. retrieving 3 of 5 facts); then define the rubric in data and grade each sub-goal programmatically so the score stays reproducible.

Key Takeaways

  1. Accuracy is one of at least four dimensions. Track completion, efficiency, cost per success, and reliability together; a single number hides expensive or brittle agents.
  2. Grade outcomes programmatically, trajectories structurally. Reserve LLM-as-judge for subjective outputs and calibrate it against human labels.
  3. An eval is reproducible data, not an anecdote. Version tasks, count crashes as failures, and re-run the identical suite to diff every change.
  4. Reliability needs repetition. Report pass@1 (the honest production number) alongside pass@k to separate variance from capability gaps.
  5. Benchmarks are coarse signals; your own suite decides. Contamination and distribution shift mean private held-out tasks drive go/no-go.
  6. Graceful failure is measurable. An agent that knows when to stop beats one that thrashes to its budget.

Next Steps: Lesson 9: Real-World Agent Patterns