Orchestration Patterns Cheat Sheet
A one-page decision aid for choosing a topology when wiring multiple agents together. Match the pattern to your control flow, fan-out, and failure tolerance, then read the production notes before you ship. Deeper treatment: course landing.
Pattern Comparison
| Pattern | Control flow | When to use | Coordination cost | Failure behaviour |
|---|---|---|---|---|
| Supervisor | Centralised — one coordinator delegates and aggregates. | Heterogeneous specialists, dynamic routing, single point of accountability. | Low–medium. One hop per delegation; coordinator is bottleneck + SPOF. | Coordinator crash halts all; worker failure is contained and retryable. |
| Peer-to-peer | Decentralised — agents message directly, no authority. | Negotiation, consensus, loosely-coupled work with no plan owner. | High. O(n²) edges; needs a shared protocol + conflict resolution. | No SPOF, but failures partition the mesh and can deadlock or loop. |
| Hierarchical | Tree of supervisors over managers over workers. | Large task trees, org-shaped domains, beyond one span of control. | Medium–high. Multiple hops; latency stacks per level. | Localised — a subtree fails alone; parent degrades or reroutes. |
| Pipeline | Sequential stages consuming the prior output. | ETL transforms, staged refinement, strict ordering. | Lowest. Linear; backpressure via the queue. | Stage failure stops the line unless buffered; per-stage retry if idempotent. |
Topologies in Text
SUPERVISOR PEER-TO-PEER
┌──────────┐ A ──── B
┌────┤coordinator├────┐ │ \ / │
▼ └────┬─────┘ ▼ │ \ / │
worker worker worker D ──C──┘ (any-to-any mesh)
HIERARCHICAL PIPELINE
┌───top───┐ in ─►[S1]─►[S2]─►[S3]─► out
┌──┴──┐ ┌──┴──┐ extract validate summarise
mgrA mgrB (each stage = queue + worker)
┌┴┐ ┌┴┐
w w w w
Selection Heuristics
| Signal | Lean toward |
|---|---|
| Skill-differentiated tasks, runtime routing | Supervisor |
| Fixed sequence transforming one artifact | Pipeline |
| > ~8 workers or nested sub-problems | Hierarchical |
| Negotiation / no clear owner | Peer-to-peer (last resort) |
| Hard latency SLO | Pipeline or shallow supervisor (fewest hops) |
| Must survive any single node | Hierarchical or P2P |
Minimal Supervisor Loop
A supervisor is a plan-act-observe loop over a worker registry. Keep delegation idempotent (stable task_id) so a retry after a crash never double-executes.
from anthropic import Anthropic
client = Anthropic()
WORKERS = {"search": search_agent, "synthesize": synth_agent, "verify": fact_agent}
def supervise(goal: str, max_steps: int = 12) -> str:
state = {"goal": goal, "results": {}, "done": False}
for step in range(max_steps):
decision = plan_next(client, state) # -> {"worker": name, "input": ...} | {"done": True}
if decision.get("done"):
state["done"] = True
break
name = decision["worker"]
task_id = f"{name}:{step}" # stable id => idempotent retry
try:
state["results"][task_id] = dispatch(name, decision["input"], task_id)
except WorkerError as e:
state["results"][task_id] = handle_failure(name, e) # retry / fallback / escalate
return aggregate(state["results"])
def dispatch(name: str, payload, task_id: str, attempts: int = 3):
for i in range(attempts):
try:
return WORKERS[name].run(payload, task_id=task_id) # idempotent worker
except TransientError:
backoff(base=0.5, attempt=i, jitter=True) # backoff + jitter
raise WorkerError(name)
plan_next is the only non-deterministic step — bound it with a token budget and a hard step cap (max_steps) so a confused coordinator cannot loop forever.
Wrapping Agents in a Workflow Engine
For durability, push the loop into a workflow engine (Temporal, Prefect, Airflow). The engine owns retries, timeouts, and the audit trail; the agent call is just an activity that survives process restarts and supports replay.
from temporalio import workflow, activity
from datetime import timedelta
@activity.defn
async def run_agent(name: str, payload: dict) -> dict:
return WORKERS[name].run(payload) # non-determinism lives in activities only
@workflow.defn
class ResearchWorkflow:
@workflow.run
async def run(self, goal: str) -> str:
hits = await workflow.execute_activity(
run_agent, args=["search", {"q": goal}],
start_to_close_timeout=timedelta(seconds=120),
retry_policy=workflow.RetryPolicy(maximum_attempts=3),
)
draft = await workflow.execute_activity(
run_agent, args=["synthesize", hits],
start_to_close_timeout=timedelta(seconds=180),
)
return draft["text"] # engine owns retries, timeouts, replay
Rule of thumb: deterministic glue in workflow code, non-determinism in activities. Never call an LLM directly from replay-based workflow code. Prefect and Airflow follow the same split.
Production Concerns by Pattern
| Concern | Supervisor | Pipeline | Hierarchical | Peer-to-peer |
|---|---|---|---|---|
| Backpressure | Coordinator throttles | Bounded queue/stage | Per-level queues | Per-edge flow |
| Tracing | One trace, fan-out spans | Stage chain | Nested per level | Hard — causal IDs |
| Cost cap | Central budget | Per-stage budget | Cascaded down tree | Per-agent, overspends |
| Consistency | Saga from coordinator | Idempotent + checkpoint | Subtree sagas | Eventual; reconcile |
Failure & Recovery Defaults
| Mechanism | Default | Notes |
|---|---|---|
| Retry | 3 attempts, exponential backoff + jitter | Transient errors only; never retry non-idempotent side effects. |
| Timeout | Per activity, < SLO budget | Separate heartbeat from start-to-close. |
| Circuit breaker | Open after N consecutive worker failures | Sheds load, stops cascades. |
| DLQ | Exhausted tasks to a queue + alert | Inspect, fix, replay. |
| Escalation | Human-in-the-loop after fallback | Hand off full context, not just the error. |
| Budget | Hard token/step cap per run | Stops runaway recursive delegation. |
Quick Anti-Patterns
- God supervisor: one coordinator holding all context — split hierarchically past ~8 workers.
- Chatty mesh: P2P where a pipeline would do — O(n²) coordination for no benefit.
- No step cap on an LLM control loop — eventual infinite loop and runaway bill.
- Non-idempotent workers behind retries — duplicate side effects on every transient failure.
- Full-context handoffs everywhere — pass summaries; reserve transcripts for debugging.
Related References
- Framework reference · Troubleshooting · Glossary · Rules
- Prerequisite: Agentic AI