Evaluating Agent Performance
Learning Outcomes
- Define a multi-dimensional metric set: completion, efficiency, cost, and reliability
- Distinguish outcome-based grading from trajectory-based grading and apply each
- Build a reusable evaluation harness that runs an agent over a labelled task set
- Interpret the benchmarks SWE-bench, GAIA, and WebArena and their limits
- Measure run-to-run reliability and detect when an agent fails to recognise it is stuck
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why accuracy alone lies about agents |
| Explain | 7 min | Outcome vs trajectory grading |
| Build | 9 min | A minimal eval harness: cost and efficiency |
| Build | 8 min | Reliability across repeated runs |
| Explain | 8 min | Benchmarks: SWE-bench, GAIA, WebArena |
| Build | 8 min | Detecting "stuck" and graceful failure |
| Wrap-up | 2 min | Key takeaways and next steps |
Before You Begin
Pre-work:
- Complete Lesson 6: Building Your First Agent — you will evaluate that agent here
- Complete Lesson 7: Agent Safety & Guardrails — resource limits feed directly into efficiency and cost metrics
Shopping List:
- Python 3.10+ and the
anthropicSDK (pip install anthropic), withANTHROPIC_API_KEYexported - The Lesson 6 agent importable as a function (goal in; result + trajectory + usage out)
- 10-20 tasks you already know the correct answer to
For a classifier, accuracy tells the whole story. For an agent it hides almost everything that matters in production. Two agents can both "solve" 70% of tasks while behaving completely differently:
| Agent A | Agent B | |
|---|---|---|
| Task completion | 70% | 70% |
| Avg cost per task | $0.04 | $0.61 |
| Run-to-run consistency | 95% | 48% |
| Recognises when stuck | Yes | No (loops to cap) |
Agent A is shippable. Agent B burns fifteen times the budget, can't repeat a success, and grinds against its cap instead of giving up. Completion told you none of that. A useful evaluation tracks at least four orthogonal dimensions — task completion, efficiency (steps vs the optimal path), cost per completed task, and reliability (same input, same outcome?) — plus a fifth, graceful failure: does the agent know when it has failed? And always hold out a test set the agent never influenced, or you measure memorisation, not capability.
There are two ways to score a run, and you usually need both. Outcome grading checks the final result against a known answer. A verifier is a function from (task, final_output) to a boolean or score:
def grade_outcome(task: dict, final_output: str) -> bool:
# Exact-match verifier; real ones run tests or compare fields.
return final_output.strip().lower() == task["expected_answer"].strip().lower()
For coding agents the verifier runs the repository's unit tests against the produced patch and returns pass/fail. That is exactly how SWE-bench grades, and why it is trusted: no judgement call.
Trajectory grading inspects how the agent got there. Walk the tool_use steps and compute signals: step count, repeated calls (len(names) - len(set(names))), whether a forbidden tool like delete_file appears. Use it when the outcome looks right but the path was unsafe, or when there is no single correct answer. For open-ended tasks, LLM-as-judge has a separate model grade the output against a rubric — it scales but is fallible, so calibrate it against human labels first.
A harness is just a loop: load tasks, run the agent, grade, aggregate. Store tasks as a flat JSON list — each with an id, goal, expected_answer, and optimal_steps (your estimate of the shortest path, the denominator for efficiency). Assume the Lesson 6 agent exposes run_agent(goal) -> AgentResult, returning final text plus trajectory and token usage.
from dataclasses import dataclass
@dataclass
class TaskResult:
completed: bool; steps: int; optimal_steps: int
in_tok: int; out_tok: int; error: str | None = None
def run_eval(tasks: list[dict]) -> list[TaskResult]:
out = []
for t in tasks:
try:
r = run_agent(t["goal"]) # your Lesson 6 agent
out.append(TaskResult(
grade_outcome(t, r.final_output), len(r.trajectory),
t["optimal_steps"], r.usage.input_tokens, r.usage.output_tokens))
except Exception as e: # a crash is a failed task
out.append(TaskResult(False, 0, t["optimal_steps"], 0, 0, str(e)))
return out
The aggregates fall out directly: completion_rate, step_efficiency (mean of optimal_steps / steps, 1.0 is perfect), crash_rate, and the key one, cost_per_success — total spend (including failures) divided by successes, so an agent that fails expensively is correctly punished. Token counts come from the Anthropic API usage block times your model's current per-million-token rates. Track cost_per_success over time, not just completion_rate.
Because agents sample tokens, the same task can pass on Monday and fail on Tuesday. Reliability measures that variance with pass@k and the stricter always-pass rate: run each task k times and look at the distribution.
from collections import defaultdict
def reliability_eval(tasks: list[dict], k: int = 5) -> dict:
outcomes = defaultdict(list) # task_id -> [bool, ...]
for t in tasks:
for _ in range(k):
r = run_agent(t["goal"])
outcomes[t["id"]].append(grade_outcome(t, r.final_output))
n = len(outcomes)
return {
"pass_at_1": sum(o[0] for o in outcomes.values()) / n,
"pass_at_k": sum(any(o) for o in outcomes.values()) / n, # any of k
"always_pass": sum(all(o) for o in outcomes.values()) / n,
}
Reading the spread diagnoses the problem: high pass@k but low pass@1 means it can solve the task unreliably (tighten prompts, lower temperature, add a verification step); all three roughly equal means it is stable (wrong-but-stable is a capability gap, not variance); low always_pass with high pass@k is the dangerous case — great in a demo, weak under one-shot load.
Know the public benchmarks before building everything yourself. Each targets a different skill and grades programmatically — no vibes.
- SWE-bench evaluates models on real-world GitHub issues: given a codebase and an issue, generate a patch, scored by running the repo's unit tests in an isolated container. SWE-bench Verified is a 500-instance subset engineers confirmed solvable.
- GAIA is 466 questions needing reasoning, multimodality, browsing, and tool use, across three difficulty levels by steps/tools needed (Level 1 few steps, up to Level 3 long sequences). Headline gap: humans ~92% vs ~15% for GPT-4 with plugins at publication.
- WebArena is a realistic, self-hostable web environment with functional sites (e-commerce, forums, GitLab, CMS), scoring agents on functional task correctness.
The worst agent behaviour is not failing — it is failing while burning your whole budget and never admitting it. A good agent recognises a dead end and stops. Measure this from the trajectory by classifying each failed run:
def diagnose_failure(result, step_cap: int) -> str:
traj = result.trajectory
last = [s["name"] for s in traj if s["type"] == "tool_use"][-4:]
if len(traj) >= step_cap:
return "hit_cap" # never stopped — worst case
if len(last) == 4 and len(set(last)) == 1:
return "looping" # same tool over and over
if result.gave_up_explicitly:
return "graceful_giveup" # best failure: it knew
return "wrong_but_terminated" # finished cleanly, just wrong
Report the distribution of these buckets. Failures dominated by graceful_giveup are far safer to deploy than ones stuck at hit_cap or looping, even at identical completion rates. This ties back to the resource limits from Lesson 7 — the step cap is your backstop, but a healthy agent rarely needs it. To push failures into that bucket, give the agent an explicit report_stuck tool to call (with a reason) when blocked.
report_stuck after two honest attempts saves money and lets a human step in early — far better than one that thrashes to the cap and returns a confident wrong answer.Questions & Answers
Key Takeaways
- Accuracy is one of at least four dimensions. Track completion, efficiency, cost per success, and reliability together; a single number hides expensive or brittle agents.
- Grade outcomes programmatically, trajectories structurally. Reserve LLM-as-judge for subjective outputs and calibrate it against human labels.
- An eval is reproducible data, not an anecdote. Version tasks, count crashes as failures, and re-run the identical suite to diff every change.
- Reliability needs repetition. Report pass@1 (the honest production number) alongside pass@k to separate variance from capability gaps.
- Benchmarks are coarse signals; your own suite decides. Contamination and distribution shift mean private held-out tasks drive go/no-go.
- Graceful failure is measurable. An agent that knows when to stop beats one that thrashes to its budget.
Next Steps: Lesson 9: Real-World Agent Patterns