Error Handling & Recovery

45 min advanced Lesson 7

Learning Outcomes

  • Classify multi-agent failure modes — transient, deterministic, partial, and cascading
  • Implement budget-aware retries with exponential backoff and jitter for non-deterministic agent calls
  • Design fallback chains that degrade gracefully instead of failing hard
  • Build circuit breakers and dead-letter queues to isolate and contain failing agents
  • Define human-in-the-loop escalation policies that hand off with full context

Lesson Plan

Segment Duration Topic
Intro 3 min Why orchestration multiplies failure surface
Explain 7 min A taxonomy of multi-agent failure modes
Demo 8 min Budget-aware retries with backoff and jitter
Demo 7 min Fallback chains and degraded mode
Demo 8 min Circuit breakers and dead-letter queues
Explain 6 min Partial failure, compensation, and idempotency
Demo 4 min Human-in-the-loop escalation
Wrap-up 2 min Key takeaways, preview observability

Before You Begin

Pre-work:

Shopping List:

  • Python 3.11+ with anthropic installed (pip install anthropic)
  • A workflow engine from Lesson 6 (Temporal or Prefect) for durable retries
  • A broker for dead-letter queues (Redis, SQS, or RabbitMQ — examples use Redis)
  • Your Anthropic API key in the environment

1 A Taxonomy of Failure Modes

In Course 04 you handled errors inside a single agent. Orchestration multiplies the surface: N agents, a network between them, a message bus, and shared state — any of which can fail independently or take each other down. Before you write a single retry, classify the failure. The recovery strategy depends on the category, not the stack trace.

Failure mode Example Right response
Transient 429, 503, socket timeout Retry with backoff
Deterministic Schema violation, 400 Do NOT retry — fix or fallback
Resource Context window or token budget exceeded Decompose/summarise, then retry
Partial 4 of 5 sub-tasks succeed Compensate or escalate the failed slice
Cascading A's timeout backs up B, then C Circuit breaker, shed load
Semantic Agent "succeeds" but output is wrong Guardrails + validation, not retries

The dangerous one is the semantic failure: the API returns 200, the agent is confident, and the answer is garbage. Retries make it worse — you burn budget re-running a call that was never going to be right. That class belongs to guardrails and evals (Agentic AI lessons 7 and 8), not the error-handling layer.

NOTE
Key Insight
Treat agents like unreliable network services, because that is what they are: high-latency, non-deterministic, occasionally-down RPC endpoints. Every distributed-systems pattern you know — timeouts, bulkheads, idempotency keys — applies directly. Blindly retrying every exception is the most common production mistake; always branch on whether the error is retryable first.

2 Budget-Aware Retries: Backoff and Jitter

The first defence against transient failure is the retry, but agent retries cost real money and seconds, so bound them by a budget, not just a count. Use exponential backoff with full jitter — without jitter, a fleet that all hits a rate limit at the same instant retries in lockstep and stampedes the API again (the "thundering herd").

import random, time, anthropic

RETRYABLE = {429, 500, 502, 503, 504}

def is_retryable(err) -> bool:
    if isinstance(err, anthropic.APIStatusError):
        return err.status_code in RETRYABLE
    return isinstance(err, (anthropic.APIConnectionError, anthropic.APITimeoutError))

def call_with_retry(fn, max_attempts=4, max_seconds=30.0):
    started = time.monotonic()
    for attempt in range(max_attempts):
        try:
            return fn()
        except Exception as err:
            last = attempt == max_attempts - 1
            if not is_retryable(err) or last or time.monotonic() - started > max_seconds:
                raise
            time.sleep(random.uniform(0, min(2 ** attempt, 8)))  # backoff + jitter

Attach the budget to the task, not the call site, so one sub-task can't retry forever and starve its siblings. Honour the API's own signal too: when a 429 carries a Retry-After header, sleep for that duration instead of your computed backoff.

TIP
Let the engine do it
If you adopted a workflow engine in Lesson 6, prefer its built-in retry policy over hand-rolled loops. Temporal activities and Prefect tasks give you durable, observable retries that survive a worker crash — your hand-rolled time.sleep does not.

3 Fallback Chains and Degraded Mode

Retries handle transient failures. When an agent is genuinely down — or keeps producing unusable output — you need a fallback: a cheaper or simpler path to a good-enough answer. Design it as an ordered chain where each rung is strictly less capable but strictly more reliable than the one above it.

def run_chain(chain: list[tuple[str, callable]], task: dict) -> dict:
    errors = []
    for name, run in chain:
        try:
            result = run(task)
            result["_served_by"] = name
            result["_degraded"] = name != chain[0][0]
            return result
        except Exception as err:
            errors.append(f"{name}: {err}")
    raise RuntimeError("all fallbacks exhausted: " + " | ".join(errors))

# best model -> faster model -> deterministic stub
research_chain = [
    ("opus_synthesis",  synthesize_with_opus),
    ("haiku_synthesis", synthesize_with_haiku),
    ("cached_template", return_last_known_good),
]

In rough order of preference, the rungs are: an alternative model (smaller/faster under capacity or latency pressure), an alternative agent (a different specialist when one worker is unhealthy), a simpler approach (one pass instead of multi-step reasoning, for repeated semantic failures), and degraded mode (a cached or partial result when all else fails). The non-negotiable rule: a degraded result must be labelled as degraded, propagated through your message schema, so a downstream agent or human can tell "verified by the fact-checker" from "best guess from a stale cache."

WARNING
Silent degradation rots trust
The worst outage is the one nobody notices. If your chain quietly serves stale cache for three hours, your dashboards stay green while users get wrong answers. Emit a metric and a structured log every time _degraded is true.

4 Circuit Breakers: Stop Hammering a Dead Agent

Retries and fallbacks operate per-request. A circuit breaker operates per-dependency: when one agent or service is clearly failing, it trips and fails fast for everyone instead of letting a thousand requests each spend 30 seconds timing out. This is your primary defence against cascading failure. A breaker has three states: closed (healthy), open (tripped, fail instantly), and half-open (let one probe through to test recovery).

import time

class CircuitOpenError(Exception): pass

class CircuitBreaker:
    def __init__(self, fail_threshold=5, reset_timeout=30.0):
        self.fail_threshold, self.reset_timeout = fail_threshold, reset_timeout
        self.failures, self.opened_at, self.state = 0, 0.0, "closed"

    def call(self, fn, *args, **kwargs):
        if self.state == "open":
            if time.monotonic() - self.opened_at < self.reset_timeout:
                raise CircuitOpenError("breaker open; failing fast")
            self.state = "half_open"                 # allow one probe
        try:
            result = fn(*args, **kwargs)
        except Exception:
            self.failures += 1
            if self.failures >= self.fail_threshold or self.state == "half_open":
                self.state, self.opened_at = "open", time.monotonic()
            raise
        self.failures, self.state = 0, "closed"      # success resets
        return result

Give every external dependency its own breaker — the Anthropic API, each MCP tool server, your vector DB. This is the bulkhead pattern: a failing fact-check tool trips its own breaker without taking down the search agent. When a breaker is open, route to the fallback chain from Step 3 rather than surfacing a raw error.

TIP
Tune to your traffic
A fail_threshold of 5 suits a low-volume coordinator. For a high-throughput worker pool, prefer a rolling error-rate (trip above 50% failures over the last 20 calls) so a few stragglers don't trip the breaker during normal operation.

5 Dead-Letter Queues: Quarantine, Don't Discard

When retries are exhausted, the breaker is open, and every fallback has failed, you have a task that cannot complete right now. Do not drop it; do not let it block the queue. Move it to a dead-letter queue (DLQ) — a quarantine where poison messages wait for inspection, manual replay, or reprocessing once the issue is fixed.

import json, time, redis
r = redis.Redis()

def dead_letter(task: dict, error: str, attempts: int):
    r.lpush("dlq:agents", json.dumps({
        "task": task, "error": error, "attempts": attempts,
        "failed_at": time.time(),
        "correlation_id": task["correlation_id"],   # stitch to its trace later
    }))
    if r.llen("dlq:agents") > 50:                    # alert on depth, not each failure
        page_on_call("DLQ depth exceeded 50 — systemic failure likely")

def replay_dlq(max_items=10):                        # operator-triggered after a fix
    for _ in range(max_items):
        raw = r.rpop("dlq:agents")
        if raw is None:
            break
        enqueue_task(json.loads(raw)["task"])        # back into the live queue

Capture the full envelope — error, attempt count, and the correlation ID so you can stitch the failure to its trace in Lesson 8's observability stack. Alert on DLQ depth, not individual entries: one dead letter is noise, fifty in five minutes is an incident. And make replay idempotent — if a task was half-completed before it died, replaying it must not double-charge a customer, which leads straight to the next step.

NOTE
Why not retry forever?
An infinitely-retried poison message is a self-inflicted DoS: it consumes a worker, fails, requeues, and starves every healthy task behind it. The DLQ converts an unbounded failure loop into a bounded, observable, operator-controlled one.

6 Partial Failure, Compensation, and Idempotency

The hardest failure in orchestration is the partial one. A supervisor fans out five sub-tasks (Lesson 4); three succeed, one is retrying, one has hit the DLQ. The work is inconsistent, and there is no database transaction spanning five LLM calls. You cannot ROLLBACK. The distributed-systems answer is the Saga pattern: for every action with a side effect, define a compensating action that undoes it, and on partial failure run them in reverse order.

class Saga:
    def __init__(self):
        self.done = []                               # (name, compensate_fn), in order

    def step(self, name, do_fn, compensate_fn):
        result = do_fn()
        self.done.append((name, compensate_fn))
        return result

    def compensate(self):                            # LIFO unwind on failure
        for name, comp in reversed(self.done):
            try:
                comp()
            except Exception as err:                 # compensation failure -> escalate
                dead_letter({"compensation": name}, str(err), attempts=1)

saga = Saga()
try:
    saga.step("extract", extract_agent.run, lambda: storage.delete(doc_id))
    saga.step("publish", lambda: publish(doc_id), lambda: unpublish(doc_id))
except Exception:
    saga.compensate()                                # undo the extract if publish fails
    raise

Compensation only works if your steps are idempotent and addressable. Idempotency keys are the single most important habit: generate a deterministic key per logical operation and have every side-effecting tool check it before acting, so a replayed "send invoice" is a no-op rather than a double charge. State checkpointing is the complement — a crashed workflow resumes from the last completed step, not from scratch.

WARNING
Eventual consistency is the default
A multi-agent system is eventually consistent, never strongly consistent. There is always a window where some sub-tasks are done and others aren't. Design downstream consumers to tolerate it — and never present a partial result as final.
TIP
Let the engine checkpoint for you
This is why Lesson 6 reached for Temporal/Prefect. A durable workflow records each completed step; on a crash it replays history and resumes, giving you saga-style recovery and checkpointing without writing the bookkeeping yourself.

7 Human-in-the-Loop Escalation

Some failures should never be auto-recovered: low confidence on an irreversible action, a compensation that itself failed, or a task that has bounced through the DLQ twice. The goal of escalation is to hand off with enough context that the human can act in seconds — not page someone who must then reconstruct the run. Define explicit triggers, and build a payload that lets the human review rather than redo.

def should_escalate(task, error, attempts) -> bool:
    return (attempts >= 3                           # exhausted recovery
            or task.get("confidence", 1.0) < 0.4    # agent self-flagged
            or task.get("irreversible", False)      # money, deletes, emails
            or isinstance(error, CompensationError))  # rollback itself failed

def escalate(task, error, history):
    notify_human({
        "summary": f"Task {task['correlation_id']} needs review",
        "what_failed": str(error),
        "attempts": history,                          # timeline, not just last error
        "proposed_action": task.get("draft_output"),  # human approves or edits
        "trace_url": f"https://traces.internal/{task['correlation_id']}",
    })
    park_task(task)    # hold, don't drop, while awaiting a decision

A good payload includes the timeline of attempts, a proposed action the human can approve or edit (not a blank prompt), and a link to the trace. Present the work the agent already did.

NOTE
Escalation is a feature, not defeat
Course 04 framed human approval as a guardrail for risky actions. In orchestration it is also a recovery path: when automation runs out of road, a clean handoff keeps the overall workflow moving instead of failing the whole job. Design it as carefully as the happy path.

Questions & Answers

Q: How do I retry a non-deterministic LLM call safely — won't a retry just produce different output?
Retry the *transport*, not the *intent*. For transient errors (429, 503, timeouts) the call never completed, so a fresh attempt is expected. The risk is side effects: if the first attempt already invoked a tool before the error, the retry could double it — guard every side-effecting tool with an idempotency key. For semantic failures where the model returns valid-but-wrong output, do not retry; that's a guardrail/eval problem.
Q: My supervisor fans out 20 workers and one fails. Fail the whole batch or return partial results?
Depends on whether the slices are independent. If they are (summarise 20 documents), return the 19 successes, DLQ the failure, and mark the result incomplete with the missing IDs. If they're interdependent (steps in one reasoning chain), one failure usually invalidates the whole result — fail fast and run compensations. The Lesson 3 schema should carry per-slice status so the supervisor can decide programmatically.
Q: Circuit breakers, DLQs, sagas — isn't this the same machinery I'd build for any microservice fleet?
Largely yes, and that's the point: the proven distributed-systems patterns transfer directly. The differences are that agent calls are far slower (seconds, not milliseconds), far more expensive per call, and can fail *semantically* while returning 200. So you tune thresholds for longer latencies, make retry budgets cost-aware, and add a validation layer pure microservices don't need.
Q: How do I keep retries from blowing my token budget during an incident?
Bind retries to a budget in attempts, seconds, and ideally tokens — track cumulative tokens per task and stop once a ceiling is hit. Pair with circuit breakers so a widespread outage trips after a handful of failures instead of letting every task burn its full retry budget. Cost-aware routing and hard caps are covered in Lesson 9.
Q: Where should this logic live — in each agent, the supervisor, or the workflow engine?
Layer it. Per-call retries belong closest to the call (or in the engine's activity retry policy). Circuit breakers wrap each external dependency. Fallback chains, sagas, and escalation are orchestration-level concerns and belong in the supervisor or workflow definition, because only that layer sees the whole task graph. Keep agents simple — they should fail loudly with structured errors and let the orchestration layer decide recovery.

Key Takeaways

  1. Classify before you recover — transient errors get retries, deterministic errors get fallbacks, semantic errors get guardrails. Retrying the wrong category burns money and latency.
  2. Bound retries by budget, not just count — backoff with jitter prevents thundering herds; attempt, time, and token ceilings stop one task from starving the fleet.
  3. Fallback chains degrade gracefully — drop to cheaper models, alternate agents, or labelled cached results, and always propagate a _degraded flag so silent degradation can't rot trust.
  4. Circuit breakers and DLQs contain blast radius — break per-dependency to stop cascades, and quarantine poison messages instead of dropping them or looping forever.
  5. Idempotency and sagas tame partial failure — there's no transaction across N agents, so make side effects idempotent and define compensating actions to unwind partial work.
  6. Escalate with context, not blame — explicit triggers, a timeline plus proposed action plus trace link, and a handoff treated as a first-class recovery path.

Next Steps: Lesson 8: Monitoring & Observability