Orchestration Patterns Cheat Sheet

Reference intermediate

A one-page decision aid for choosing a topology when wiring multiple agents together. Match the pattern to your control flow, fan-out, and failure tolerance, then read the production notes before you ship. Deeper treatment: course landing.

Pattern Comparison

Pattern Control flow When to use Coordination cost Failure behaviour
Supervisor Centralised — one coordinator delegates and aggregates. Heterogeneous specialists, dynamic routing, single point of accountability. Low–medium. One hop per delegation; coordinator is bottleneck + SPOF. Coordinator crash halts all; worker failure is contained and retryable.
Peer-to-peer Decentralised — agents message directly, no authority. Negotiation, consensus, loosely-coupled work with no plan owner. High. O(n²) edges; needs a shared protocol + conflict resolution. No SPOF, but failures partition the mesh and can deadlock or loop.
Hierarchical Tree of supervisors over managers over workers. Large task trees, org-shaped domains, beyond one span of control. Medium–high. Multiple hops; latency stacks per level. Localised — a subtree fails alone; parent degrades or reroutes.
Pipeline Sequential stages consuming the prior output. ETL transforms, staged refinement, strict ordering. Lowest. Linear; backpressure via the queue. Stage failure stops the line unless buffered; per-stage retry if idempotent.

Topologies in Text

SUPERVISOR                  PEER-TO-PEER
        ┌──────────┐                A ──── B
   ┌────┤coordinator├────┐          │ \   / │
   ▼    └────┬─────┘    ▼          │  \ /  │
worker     worker     worker       D ──C──┘  (any-to-any mesh)

HIERARCHICAL                PIPELINE
        ┌───top───┐          in ─►[S1]─►[S2]─►[S3]─► out
     ┌──┴──┐   ┌──┴──┐         extract validate summarise
   mgrA        mgrB           (each stage = queue + worker)
   ┌┴┐         ┌┴┐
  w  w        w  w

Selection Heuristics

Signal Lean toward
Skill-differentiated tasks, runtime routing Supervisor
Fixed sequence transforming one artifact Pipeline
> ~8 workers or nested sub-problems Hierarchical
Negotiation / no clear owner Peer-to-peer (last resort)
Hard latency SLO Pipeline or shallow supervisor (fewest hops)
Must survive any single node Hierarchical or P2P

Minimal Supervisor Loop

A supervisor is a plan-act-observe loop over a worker registry. Keep delegation idempotent (stable task_id) so a retry after a crash never double-executes.

from anthropic import Anthropic

client = Anthropic()
WORKERS = {"search": search_agent, "synthesize": synth_agent, "verify": fact_agent}

def supervise(goal: str, max_steps: int = 12) -> str:
    state = {"goal": goal, "results": {}, "done": False}
    for step in range(max_steps):
        decision = plan_next(client, state)          # -> {"worker": name, "input": ...} | {"done": True}
        if decision.get("done"):
            state["done"] = True
            break
        name = decision["worker"]
        task_id = f"{name}:{step}"                    # stable id => idempotent retry
        try:
            state["results"][task_id] = dispatch(name, decision["input"], task_id)
        except WorkerError as e:
            state["results"][task_id] = handle_failure(name, e)   # retry / fallback / escalate
    return aggregate(state["results"])

def dispatch(name: str, payload, task_id: str, attempts: int = 3):
    for i in range(attempts):
        try:
            return WORKERS[name].run(payload, task_id=task_id)   # idempotent worker
        except TransientError:
            backoff(base=0.5, attempt=i, jitter=True)            # backoff + jitter
    raise WorkerError(name)

plan_next is the only non-deterministic step — bound it with a token budget and a hard step cap (max_steps) so a confused coordinator cannot loop forever.

Wrapping Agents in a Workflow Engine

For durability, push the loop into a workflow engine (Temporal, Prefect, Airflow). The engine owns retries, timeouts, and the audit trail; the agent call is just an activity that survives process restarts and supports replay.

from temporalio import workflow, activity
from datetime import timedelta

@activity.defn
async def run_agent(name: str, payload: dict) -> dict:
    return WORKERS[name].run(payload)        # non-determinism lives in activities only

@workflow.defn
class ResearchWorkflow:
    @workflow.run
    async def run(self, goal: str) -> str:
        hits = await workflow.execute_activity(
            run_agent, args=["search", {"q": goal}],
            start_to_close_timeout=timedelta(seconds=120),
            retry_policy=workflow.RetryPolicy(maximum_attempts=3),
        )
        draft = await workflow.execute_activity(
            run_agent, args=["synthesize", hits],
            start_to_close_timeout=timedelta(seconds=180),
        )
        return draft["text"]                 # engine owns retries, timeouts, replay

Rule of thumb: deterministic glue in workflow code, non-determinism in activities. Never call an LLM directly from replay-based workflow code. Prefect and Airflow follow the same split.

Production Concerns by Pattern

Concern Supervisor Pipeline Hierarchical Peer-to-peer
Backpressure Coordinator throttles Bounded queue/stage Per-level queues Per-edge flow
Tracing One trace, fan-out spans Stage chain Nested per level Hard — causal IDs
Cost cap Central budget Per-stage budget Cascaded down tree Per-agent, overspends
Consistency Saga from coordinator Idempotent + checkpoint Subtree sagas Eventual; reconcile

Failure & Recovery Defaults

Mechanism Default Notes
Retry 3 attempts, exponential backoff + jitter Transient errors only; never retry non-idempotent side effects.
Timeout Per activity, < SLO budget Separate heartbeat from start-to-close.
Circuit breaker Open after N consecutive worker failures Sheds load, stops cascades.
DLQ Exhausted tasks to a queue + alert Inspect, fix, replay.
Escalation Human-in-the-loop after fallback Hand off full context, not just the error.
Budget Hard token/step cap per run Stops runaway recursive delegation.

Quick Anti-Patterns

  • God supervisor: one coordinator holding all context — split hierarchically past ~8 workers.
  • Chatty mesh: P2P where a pipeline would do — O(n²) coordination for no benefit.
  • No step cap on an LLM control loop — eventual infinite loop and runaway bill.
  • Non-idempotent workers behind retries — duplicate side effects on every transient failure.
  • Full-context handoffs everywhere — pass summaries; reserve transcripts for debugging.

Related References