Troubleshooting

Reference intermediate

Multi-agent failures rarely arrive with a clean stack trace. They surface as a workflow that hangs, a bill that triples overnight, or a "completed" run that produced nothing. This page maps common production failure modes to root causes and fixes. Pair it with the rules and the framework reference.

Triage: symptom to suspect

Symptom Likely failure class
Whole pipeline errors after one agent fails Cascading failure
Run reports success but some outputs missing Partial failure
Agent acts on stale or out-of-order context Message-ordering bug
Workflow hangs with no progress, no error Deadlock
Agents keep re-triggering each other, no progress Livelock
Token spend spikes with no extra output Runaway cost
Final result is empty / dropped Lost results

Cascading failures

One agent throws, every downstream agent treats its output as a hard dependency, and the whole DAG collapses. Supervisor retries then amplify load on a degraded sub-agent (the retry storm).

Problem Cause Fix
Single agent error kills the run No fault isolation; supervisor awaits all children Circuit-break each call; treat a tripped breaker as a typed result
Retry storm overloads a slow agent Synchronous retries, no budget Backoff with full jitter; cap attempts per workflow
Optional agent's failure aborts the goal No criticality tier Tag critical vs best-effort; only critical failures abort

In a workflow engine, give each activity its own start_to_close_timeout and a RetryPolicy (maximum_attempts, backoff_coefficient), and mark guardrail violations non-retryable so they fail fast.

Partial failures

In fan-out work, three of five agents succeed and two error. Default gather semantics let the first exception discard the good work too.

Problem Cause Fix
Good sub-results lost on first failure asyncio.gather raises on first error gather(..., return_exceptions=True); keep survivors
Reassembly assumes all parts present Aggregator indexes by position Key by task id; reconcile vs expected set
Cannot retry just the failed shard No per-task checkpoint Persist per-shard state; re-dispatch failed only

Decide the aggregation contract up front — all shards, a quorum, or best-effort — so the orchestrator knows whether failed means "abort" or "ship what we have."

Message-ordering bugs

Multi-agent systems are distributed systems. Messages arrive out of order, get redelivered, or interleave between supervisors, and an agent reasons over a context that never existed.

Problem Cause Fix
Agent reads stale shared state Read-modify-write race on the blackboard Optimistic concurrency: version state, reject stale writes
Same event processed twice At-least-once delivery (default) Idempotency keys; dedupe by (agent, event_id)
Replies attributed to wrong request No correlation id across hops Stamp trace_id + causation_id on every message
Out-of-order tool results Sequential assumptions on parallel calls Buffer to sequence, or make handlers order-free

Prefer event-sourced state: append immutable facts keyed by causation and let each agent fold the log into its own view, backing the dedupe set with Redis so idempotency survives a restart. This trades ordering races for eventual consistency.

Deadlocks and livelocks

Two agents each wait on the other (deadlock), or hand the same task back and forth without converging (livelock). Peer-to-peer and hierarchical topologies are the usual culprits:

Problem Cause Fix
Two agents block on each other Circular dependency in the DAG Detect cycles at plan time; forbid the edge or add a tie-break
Workflow hangs forever, no error Missing timeout on a handoff Wrap every await in a timeout; return a typed timeout result
Reviewer/coder loop never ends No convergence criterion or turn cap Cap handoff rounds; escalate to human on exhaustion
Shared lock never released Agent crashed holding a lease Use TTL leases, not locks; auto-expire on owner death

Wrap every handoff await in asyncio.wait_for returning a typed timeout so a stuck agent surfaces instead of hanging. For livelock, use a monotonic progress measure: if a loop does not reduce some objective within N rounds, break it.

Runaway token cost

The most common production surprise. A retry loop, an over-eager re-planner, or an agent re-sending the whole conversation each turn multiplies spend silently because nothing crashes.

Problem Cause Fix
Spend spikes, no extra output Retry loop re-sending full context Budget-aware retries; abort when budget exhausted
Every agent gets the whole transcript Passing full context, not summaries Hand off structured state or summaries, not raw history
Same sub-task computed repeatedly No result cache Cache by content hash of the sub-task input
Identical system prompt re-billed No prompt caching across calls Cache stable prefixes (instructions, tool defs)
Simple tasks hit the largest model No routing tier Route by complexity: small model triages, large for hard steps

Scope a shared TokenBudget (accumulates used, raises past cap) per workflow run, not per call — per-call limits miss a loop of 500 cheap calls. Emit used/cap and alert at 80 percent. See the course for routing and caching.

Lost results

The run completes, logs look clean, and the final payload is empty or truncated — the hardest class to debug, with no error to grep for.

Problem Cause Fix
Final output empty after success Fire-and-forget call not awaited Await every spawned task; never create_task without joining
Result silently truncated Output exceeded model max tokens Check stop_reason; if max_tokens, continue or chunk
Aggregator overwrites prior writes Concurrent writes to one shared key Append to keyed slots, not a shared scalar; reconcile at end
Worker output never persisted Died after compute, before commit Persist before acking the message (process-then-ack)
Dead-lettered work never retried DLQ has no consumer Wire a DLQ drain job with an escalation policy

Use process-then-ack: durably store.put the result before queue.ack so a crash redelivers the work. And always assert on stop_reason — a response that stopped at the output limit is not complete, and treating it as one is the top cause of truncation.

When triaging any of these, gather evidence first: trace where the run stopped via trace_id; compare per-activity latency against timeouts; count invocations per event_id to detect duplication; and diff persisted store keys against the expected set.

Related references