Error Handling & Recovery
Learning Outcomes
- Classify multi-agent failure modes — transient, deterministic, partial, and cascading
- Implement budget-aware retries with exponential backoff and jitter for non-deterministic agent calls
- Design fallback chains that degrade gracefully instead of failing hard
- Build circuit breakers and dead-letter queues to isolate and contain failing agents
- Define human-in-the-loop escalation policies that hand off with full context
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why orchestration multiplies failure surface |
| Explain | 7 min | A taxonomy of multi-agent failure modes |
| Demo | 8 min | Budget-aware retries with backoff and jitter |
| Demo | 7 min | Fallback chains and degraded mode |
| Demo | 8 min | Circuit breakers and dead-letter queues |
| Explain | 6 min | Partial failure, compensation, and idempotency |
| Demo | 4 min | Human-in-the-loop escalation |
| Wrap-up | 2 min | Key takeaways, preview observability |
Before You Begin
Pre-work:
- Complete Lesson 6: Workflow Engines — we build on durable retries here
- Review Agent Safety & Guardrails — guardrails catch bad output; this lesson catches failed output
- Have a working multi-agent system from earlier lessons (coordinator + workers)
Shopping List:
- Python 3.11+ with
anthropicinstalled (pip install anthropic) - A workflow engine from Lesson 6 (Temporal or Prefect) for durable retries
- A broker for dead-letter queues (Redis, SQS, or RabbitMQ — examples use Redis)
- Your Anthropic API key in the environment
In Course 04 you handled errors inside a single agent. Orchestration multiplies the surface: N agents, a network between them, a message bus, and shared state — any of which can fail independently or take each other down. Before you write a single retry, classify the failure. The recovery strategy depends on the category, not the stack trace.
| Failure mode | Example | Right response |
|---|---|---|
| Transient | 429, 503, socket timeout | Retry with backoff |
| Deterministic | Schema violation, 400 | Do NOT retry — fix or fallback |
| Resource | Context window or token budget exceeded | Decompose/summarise, then retry |
| Partial | 4 of 5 sub-tasks succeed | Compensate or escalate the failed slice |
| Cascading | A's timeout backs up B, then C | Circuit breaker, shed load |
| Semantic | Agent "succeeds" but output is wrong | Guardrails + validation, not retries |
The dangerous one is the semantic failure: the API returns 200, the agent is confident, and the answer is garbage. Retries make it worse — you burn budget re-running a call that was never going to be right. That class belongs to guardrails and evals (Agentic AI lessons 7 and 8), not the error-handling layer.
retryable first.The first defence against transient failure is the retry, but agent retries cost real money and seconds, so bound them by a budget, not just a count. Use exponential backoff with full jitter — without jitter, a fleet that all hits a rate limit at the same instant retries in lockstep and stampedes the API again (the "thundering herd").
import random, time, anthropic
RETRYABLE = {429, 500, 502, 503, 504}
def is_retryable(err) -> bool:
if isinstance(err, anthropic.APIStatusError):
return err.status_code in RETRYABLE
return isinstance(err, (anthropic.APIConnectionError, anthropic.APITimeoutError))
def call_with_retry(fn, max_attempts=4, max_seconds=30.0):
started = time.monotonic()
for attempt in range(max_attempts):
try:
return fn()
except Exception as err:
last = attempt == max_attempts - 1
if not is_retryable(err) or last or time.monotonic() - started > max_seconds:
raise
time.sleep(random.uniform(0, min(2 ** attempt, 8))) # backoff + jitter
Attach the budget to the task, not the call site, so one sub-task can't retry forever and starve its siblings. Honour the API's own signal too: when a 429 carries a Retry-After header, sleep for that duration instead of your computed backoff.
time.sleep does not.Retries handle transient failures. When an agent is genuinely down — or keeps producing unusable output — you need a fallback: a cheaper or simpler path to a good-enough answer. Design it as an ordered chain where each rung is strictly less capable but strictly more reliable than the one above it.
def run_chain(chain: list[tuple[str, callable]], task: dict) -> dict:
errors = []
for name, run in chain:
try:
result = run(task)
result["_served_by"] = name
result["_degraded"] = name != chain[0][0]
return result
except Exception as err:
errors.append(f"{name}: {err}")
raise RuntimeError("all fallbacks exhausted: " + " | ".join(errors))
# best model -> faster model -> deterministic stub
research_chain = [
("opus_synthesis", synthesize_with_opus),
("haiku_synthesis", synthesize_with_haiku),
("cached_template", return_last_known_good),
]
In rough order of preference, the rungs are: an alternative model (smaller/faster under capacity or latency pressure), an alternative agent (a different specialist when one worker is unhealthy), a simpler approach (one pass instead of multi-step reasoning, for repeated semantic failures), and degraded mode (a cached or partial result when all else fails). The non-negotiable rule: a degraded result must be labelled as degraded, propagated through your message schema, so a downstream agent or human can tell "verified by the fact-checker" from "best guess from a stale cache."
_degraded is true.Retries and fallbacks operate per-request. A circuit breaker operates per-dependency: when one agent or service is clearly failing, it trips and fails fast for everyone instead of letting a thousand requests each spend 30 seconds timing out. This is your primary defence against cascading failure. A breaker has three states: closed (healthy), open (tripped, fail instantly), and half-open (let one probe through to test recovery).
import time
class CircuitOpenError(Exception): pass
class CircuitBreaker:
def __init__(self, fail_threshold=5, reset_timeout=30.0):
self.fail_threshold, self.reset_timeout = fail_threshold, reset_timeout
self.failures, self.opened_at, self.state = 0, 0.0, "closed"
def call(self, fn, *args, **kwargs):
if self.state == "open":
if time.monotonic() - self.opened_at < self.reset_timeout:
raise CircuitOpenError("breaker open; failing fast")
self.state = "half_open" # allow one probe
try:
result = fn(*args, **kwargs)
except Exception:
self.failures += 1
if self.failures >= self.fail_threshold or self.state == "half_open":
self.state, self.opened_at = "open", time.monotonic()
raise
self.failures, self.state = 0, "closed" # success resets
return result
Give every external dependency its own breaker — the Anthropic API, each MCP tool server, your vector DB. This is the bulkhead pattern: a failing fact-check tool trips its own breaker without taking down the search agent. When a breaker is open, route to the fallback chain from Step 3 rather than surfacing a raw error.
fail_threshold of 5 suits a low-volume coordinator. For a high-throughput worker pool, prefer a rolling error-rate (trip above 50% failures over the last 20 calls) so a few stragglers don't trip the breaker during normal operation.When retries are exhausted, the breaker is open, and every fallback has failed, you have a task that cannot complete right now. Do not drop it; do not let it block the queue. Move it to a dead-letter queue (DLQ) — a quarantine where poison messages wait for inspection, manual replay, or reprocessing once the issue is fixed.
import json, time, redis
r = redis.Redis()
def dead_letter(task: dict, error: str, attempts: int):
r.lpush("dlq:agents", json.dumps({
"task": task, "error": error, "attempts": attempts,
"failed_at": time.time(),
"correlation_id": task["correlation_id"], # stitch to its trace later
}))
if r.llen("dlq:agents") > 50: # alert on depth, not each failure
page_on_call("DLQ depth exceeded 50 — systemic failure likely")
def replay_dlq(max_items=10): # operator-triggered after a fix
for _ in range(max_items):
raw = r.rpop("dlq:agents")
if raw is None:
break
enqueue_task(json.loads(raw)["task"]) # back into the live queue
Capture the full envelope — error, attempt count, and the correlation ID so you can stitch the failure to its trace in Lesson 8's observability stack. Alert on DLQ depth, not individual entries: one dead letter is noise, fifty in five minutes is an incident. And make replay idempotent — if a task was half-completed before it died, replaying it must not double-charge a customer, which leads straight to the next step.
The hardest failure in orchestration is the partial one. A supervisor fans out five sub-tasks (Lesson 4); three succeed, one is retrying, one has hit the DLQ. The work is inconsistent, and there is no database transaction spanning five LLM calls. You cannot ROLLBACK. The distributed-systems answer is the Saga pattern: for every action with a side effect, define a compensating action that undoes it, and on partial failure run them in reverse order.
class Saga:
def __init__(self):
self.done = [] # (name, compensate_fn), in order
def step(self, name, do_fn, compensate_fn):
result = do_fn()
self.done.append((name, compensate_fn))
return result
def compensate(self): # LIFO unwind on failure
for name, comp in reversed(self.done):
try:
comp()
except Exception as err: # compensation failure -> escalate
dead_letter({"compensation": name}, str(err), attempts=1)
saga = Saga()
try:
saga.step("extract", extract_agent.run, lambda: storage.delete(doc_id))
saga.step("publish", lambda: publish(doc_id), lambda: unpublish(doc_id))
except Exception:
saga.compensate() # undo the extract if publish fails
raise
Compensation only works if your steps are idempotent and addressable. Idempotency keys are the single most important habit: generate a deterministic key per logical operation and have every side-effecting tool check it before acting, so a replayed "send invoice" is a no-op rather than a double charge. State checkpointing is the complement — a crashed workflow resumes from the last completed step, not from scratch.
Some failures should never be auto-recovered: low confidence on an irreversible action, a compensation that itself failed, or a task that has bounced through the DLQ twice. The goal of escalation is to hand off with enough context that the human can act in seconds — not page someone who must then reconstruct the run. Define explicit triggers, and build a payload that lets the human review rather than redo.
def should_escalate(task, error, attempts) -> bool:
return (attempts >= 3 # exhausted recovery
or task.get("confidence", 1.0) < 0.4 # agent self-flagged
or task.get("irreversible", False) # money, deletes, emails
or isinstance(error, CompensationError)) # rollback itself failed
def escalate(task, error, history):
notify_human({
"summary": f"Task {task['correlation_id']} needs review",
"what_failed": str(error),
"attempts": history, # timeline, not just last error
"proposed_action": task.get("draft_output"), # human approves or edits
"trace_url": f"https://traces.internal/{task['correlation_id']}",
})
park_task(task) # hold, don't drop, while awaiting a decision
A good payload includes the timeline of attempts, a proposed action the human can approve or edit (not a blank prompt), and a link to the trace. Present the work the agent already did.
Questions & Answers
Key Takeaways
- Classify before you recover — transient errors get retries, deterministic errors get fallbacks, semantic errors get guardrails. Retrying the wrong category burns money and latency.
- Bound retries by budget, not just count — backoff with jitter prevents thundering herds; attempt, time, and token ceilings stop one task from starving the fleet.
- Fallback chains degrade gracefully — drop to cheaper models, alternate agents, or labelled cached results, and always propagate a
_degradedflag so silent degradation can't rot trust. - Circuit breakers and DLQs contain blast radius — break per-dependency to stop cascades, and quarantine poison messages instead of dropping them or looping forever.
- Idempotency and sagas tame partial failure — there's no transaction across N agents, so make side effects idempotent and define compensating actions to unwind partial work.
- Escalate with context, not blame — explicit triggers, a timeline plus proposed action plus trace link, and a handoff treated as a first-class recovery path.
Next Steps: Lesson 8: Monitoring & Observability