Troubleshooting
Multi-agent failures rarely arrive with a clean stack trace. They surface as a workflow that hangs, a bill that triples overnight, or a "completed" run that produced nothing. This page maps common production failure modes to root causes and fixes. Pair it with the rules and the framework reference.
Triage: symptom to suspect
| Symptom | Likely failure class |
|---|---|
| Whole pipeline errors after one agent fails | Cascading failure |
| Run reports success but some outputs missing | Partial failure |
| Agent acts on stale or out-of-order context | Message-ordering bug |
| Workflow hangs with no progress, no error | Deadlock |
| Agents keep re-triggering each other, no progress | Livelock |
| Token spend spikes with no extra output | Runaway cost |
| Final result is empty / dropped | Lost results |
Cascading failures
One agent throws, every downstream agent treats its output as a hard dependency, and the whole DAG collapses. Supervisor retries then amplify load on a degraded sub-agent (the retry storm).
| Problem | Cause | Fix |
|---|---|---|
| Single agent error kills the run | No fault isolation; supervisor awaits all children |
Circuit-break each call; treat a tripped breaker as a typed result |
| Retry storm overloads a slow agent | Synchronous retries, no budget | Backoff with full jitter; cap attempts per workflow |
| Optional agent's failure aborts the goal | No criticality tier | Tag critical vs best-effort; only critical failures abort |
In a workflow engine, give each activity its own start_to_close_timeout and a RetryPolicy (maximum_attempts, backoff_coefficient), and mark guardrail violations non-retryable so they fail fast.
Partial failures
In fan-out work, three of five agents succeed and two error. Default gather semantics let the first exception discard the good work too.
| Problem | Cause | Fix |
|---|---|---|
| Good sub-results lost on first failure | asyncio.gather raises on first error |
gather(..., return_exceptions=True); keep survivors |
| Reassembly assumes all parts present | Aggregator indexes by position | Key by task id; reconcile vs expected set |
| Cannot retry just the failed shard | No per-task checkpoint | Persist per-shard state; re-dispatch failed only |
Decide the aggregation contract up front — all shards, a quorum, or best-effort — so the orchestrator knows whether failed means "abort" or "ship what we have."
Message-ordering bugs
Multi-agent systems are distributed systems. Messages arrive out of order, get redelivered, or interleave between supervisors, and an agent reasons over a context that never existed.
| Problem | Cause | Fix |
|---|---|---|
| Agent reads stale shared state | Read-modify-write race on the blackboard | Optimistic concurrency: version state, reject stale writes |
| Same event processed twice | At-least-once delivery (default) | Idempotency keys; dedupe by (agent, event_id) |
| Replies attributed to wrong request | No correlation id across hops | Stamp trace_id + causation_id on every message |
| Out-of-order tool results | Sequential assumptions on parallel calls | Buffer to sequence, or make handlers order-free |
Prefer event-sourced state: append immutable facts keyed by causation and let each agent fold the log into its own view, backing the dedupe set with Redis so idempotency survives a restart. This trades ordering races for eventual consistency.
Deadlocks and livelocks
Two agents each wait on the other (deadlock), or hand the same task back and forth without converging (livelock). Peer-to-peer and hierarchical topologies are the usual culprits:
| Problem | Cause | Fix |
|---|---|---|
| Two agents block on each other | Circular dependency in the DAG | Detect cycles at plan time; forbid the edge or add a tie-break |
| Workflow hangs forever, no error | Missing timeout on a handoff | Wrap every await in a timeout; return a typed timeout result |
| Reviewer/coder loop never ends | No convergence criterion or turn cap | Cap handoff rounds; escalate to human on exhaustion |
| Shared lock never released | Agent crashed holding a lease | Use TTL leases, not locks; auto-expire on owner death |
Wrap every handoff await in asyncio.wait_for returning a typed timeout so a stuck agent surfaces instead of hanging. For livelock, use a monotonic progress measure: if a loop does not reduce some objective within N rounds, break it.
Runaway token cost
The most common production surprise. A retry loop, an over-eager re-planner, or an agent re-sending the whole conversation each turn multiplies spend silently because nothing crashes.
| Problem | Cause | Fix |
|---|---|---|
| Spend spikes, no extra output | Retry loop re-sending full context | Budget-aware retries; abort when budget exhausted |
| Every agent gets the whole transcript | Passing full context, not summaries | Hand off structured state or summaries, not raw history |
| Same sub-task computed repeatedly | No result cache | Cache by content hash of the sub-task input |
| Identical system prompt re-billed | No prompt caching across calls | Cache stable prefixes (instructions, tool defs) |
| Simple tasks hit the largest model | No routing tier | Route by complexity: small model triages, large for hard steps |
Scope a shared TokenBudget (accumulates used, raises past cap) per workflow run, not per call — per-call limits miss a loop of 500 cheap calls. Emit used/cap and alert at 80 percent. See the course for routing and caching.
Lost results
The run completes, logs look clean, and the final payload is empty or truncated — the hardest class to debug, with no error to grep for.
| Problem | Cause | Fix |
|---|---|---|
| Final output empty after success | Fire-and-forget call not awaited | Await every spawned task; never create_task without joining |
| Result silently truncated | Output exceeded model max tokens | Check stop_reason; if max_tokens, continue or chunk |
| Aggregator overwrites prior writes | Concurrent writes to one shared key | Append to keyed slots, not a shared scalar; reconcile at end |
| Worker output never persisted | Died after compute, before commit | Persist before acking the message (process-then-ack) |
| Dead-lettered work never retried | DLQ has no consumer | Wire a DLQ drain job with an escalation policy |
Use process-then-ack: durably store.put the result before queue.ack so a crash redelivers the work. And always assert on stop_reason — a response that stopped at the output limit is not complete, and treating it as one is the top cause of truncation.
When triaging any of these, gather evidence first: trace where the run stopped via trace_id; compare per-activity latency against timeouts; count invocations per event_id to detect duplication; and diff persisted store keys against the expected set.
Related references
- Rules — constraints that prevent these failures
- Framework reference — SDK and workflow-engine APIs
- Cheat sheet — command and pattern lookup
- Glossary — idempotency, livelock, DLQ, and more
- Agentic AI — single-agent foundations