Orchestration Patterns

40 min intermediate Lesson 2

Learning Outcomes

  • Distinguish the four dominant orchestration patterns and their control-flow topologies
  • Map each pattern's strengths and failure modes onto distributed-systems concerns you know
  • Implement a minimal supervisor and a minimal pipeline in Python
  • Evaluate latency, cost, and fault-tolerance trade-offs before committing to a topology
  • Select the right pattern for a workload using a repeatable decision procedure

Lesson Plan

Segment Duration Topic
Intro 3 min Topology as the first architectural decision
Supervisor 8 min Central coordinator delegating to workers
P2P & hierarchical 10 min Mesh collaboration and recursive supervisors
Pipeline 7 min Sequential staged processing
Trade-offs 6 min Latency, cost, and fault tolerance compared
Selection 4 min A decision procedure with worked examples
Wrap-up 2 min Key takeaways, preview communication

Before You Begin

Pre-work:

Shopping List:

  • Python 3.11+ with pip install anthropic and an ANTHROPIC_API_KEY exported
  • A scratch project directory and your editor of choice

1 The Supervisor Pattern — Central Coordination

A supervisor (orchestrator, router) is a single coordinating agent that owns the task graph: it takes the goal, decides which specialist acts next, dispatches a sub-task, collects the result, and repeats. Workers never talk to each other — every edge passes through the supervisor, a star topology with stateless workers at the points.

This is the workhorse pattern and where most teams should start: centralized routing makes the supervisor the single source of truth, so the system is easy to reason about, log, and debug. A minimal supervisor, where worker(system, task, model) is a one-line wrapper around client.messages.create(...) returning the text:

FAST, SMART = "claude-haiku-latest", "claude-sonnet-latest"

WORKERS = {  # stateless specialists keyed by role
    "search":   lambda t: worker("Find relevant facts. Be terse.", t, FAST),
    "writer":   lambda t: worker("Draft clear prose.", t, SMART),
    "reviewer": lambda t: worker("Critique drafts for accuracy.", t, SMART)}

def supervisor(goal: str) -> str:                # route hard-coded for clarity
    draft = WORKERS["writer"](goal + " / " + WORKERS["search"](goal))
    return WORKERS["writer"]("Revise per: " + WORKERS["reviewer"](draft))

In production that route is itself a model call returning a structured action naming the next worker — made dynamic in Lesson 4: Task Decomposition at Scale.

WARNING
The supervisor centralizes both control and risk
Centralizing state puts retries, idempotency, and audit logging in one place — but every task touches the supervisor, so its context window and availability cap the system and make it your SPOF. Keep its prompt lean (route, do not reason) and scale it as its own tier.

2 The Peer-to-Peer Pattern — Direct Collaboration

In peer-to-peer, agents communicate directly with no central router; each decides who to address next from the conversation so far. This is the model behind frameworks like AutoGen — a "group chat" of specialists contributing until a termination condition fires.

   [analyst] ◀──▶ [critic]      mesh: any agent may
       ▲             ▲           address any other
       └──[coder]────┘

Emergent collaboration is the strength — agents debate and correct each other, reaching solutions a fixed pipeline never would. The cost is a non-deterministic topology that can loop, so you must bound it with a hard turn cap and an explicit TASK_COMPLETE token. The loop iterates range(max_turns), picks a speaker (round-robin or an LLM moderator), and appends each reply to a shared transcript. Without that cap and token, two agents ping-pong forever — treat max_turns as an SLO guardrail, not a tuning knob.

WARNING
P2P erases your audit trail
With no central coordinator, nothing knows the full state — you inherit message ordering, eventual consistency, and deadlock detection without a clean log. Prefer a moderated chat with one selector picking the next speaker; reach for full P2P only when emergent debate is a genuine requirement.

3 The Hierarchical Pattern — Supervisors of Supervisors

Hierarchical orchestration is the supervisor pattern applied recursively: a top-level supervisor delegates to mid-level supervisors, each owning a team of workers. It is the org-chart of agent systems — keeping any one coordinator's span of control, here its context window, manageable.

              [ EXECUTIVE ]                  tree topology
               │         │
      [RESEARCH lead]  [WRITING lead]        team leads
       │      │          │     │
    search summarize   draft  edit           workers

The advantage over a flat supervisor is context isolation: the executive holds only the goal plus per-team summaries, each lead only its objective plus its workers' results, each worker only a single task. Every prompt stays small even when the workflow is large. The trade-off is depth-proportional latency — a leaf result bubbles all the way up, each layer adding a round-trip, so a three-layer hierarchy means roughly six sequential LLM calls before the executive responds.

NOTE
When the flat supervisor stops scaling
Add a layer when the routing prompt outgrows a comfortable fraction of the context window, or when distinct sub-domains each need their own tools and policies. Below that, a flat supervisor is simpler and faster.

4 The Pipeline Pattern — Sequential Staged Processing

A pipeline chains agents into fixed stages (extract ─▶ validate ─▶ transform ─▶ summarize), the output of stage N feeding stage N+1. No router makes dynamic decisions; the DAG is baked in at design time, the multi-agent analogue of a Unix pipe or ETL job.

Pipelines win when the work has a stable shape — document processing, data enrichment, content generation with QA. The runner is trivial: fold the payload through a list of named stages, wrapping each so a failure reports which stage broke.

def pipeline(stages: list[tuple[str, callable]], payload: str) -> str:
    for name, run in stages:
        try:
            payload = run(payload)
        except Exception as e:
            raise RuntimeError("stage '" + name + "' failed: " + str(e)) from e
    return payload

Fixed stage positions make pipelines the easiest pattern to test, cache, and parallelize. The DAG also maps onto a workflow engine — each stage a durable activity with its own retry and timeout, a Prefect @task with retries=3 giving backoff for free, the subject of Lesson 6: Workflow Engines.

WARNING
Rigid by design
A pipeline cannot adapt its route mid-run; a stage needing an unanticipated branch can only fail or pass bad data forward. Throughput comes from fan-out over inputs, but per-item latency is the sum of stages — when routes must be decided at runtime, use a supervisor.

5 Trade-offs Side by Side

The four patterns sit on a spectrum from centralized-deterministic (pipeline) to decentralized-emergent (peer-to-peer). Read this as a set of dials, not a leaderboard.

Dimension Supervisor Peer-to-peer Hierarchical Pipeline
Control flow Star Mesh Tree Linear DAG
Latency 2 hops/task Unbounded* Depth × hops Sum of stages
SPOF Supervisor None Subtree roots Any stage
Observability Easy Hard Medium Easy
Token cost Medium High Med-high Predictable
Best for Dynamic routing Open-ended debate Large decomposable work Stable repeatable flows

* Bounded only by your hard max_turns guardrail.

Two themes recur: determinism and adaptivity trade off, and centralization concentrates failure but simplifies recovery — a supervisor is a SPOF, but the single home for a circuit breaker and trace (Lesson 7), while cost is itself a topology decision (Lesson 9: Cost Management & Optimization).

NOTE
Patterns compose
Real systems mix patterns: a pipeline whose middle stage is internally a supervisor, a hierarchy whose leaf teams run as pipelines. Choosing a pattern is choosing the topology for one boundary, not the whole system.

6 A Decision Procedure

For a new workload, walk these questions in order; the first "yes" is usually your answer.

Q1 Steps fixed and known in advance?                  ─▶ PIPELINE
Q2 Routing depends on runtime results?                ─▶ SUPERVISOR
Q3 Supervisor prompt/tools too large, or sub-domains
   needing distinct policies?                          ─▶ HIERARCHICAL
Q4 Task truly requires agents to debate / critique?    ─▶ PEER-TO-PEER (moderated)

Applied to real workloads: nightly invoice extract + validate + post is fixed steps (Q1 → pipeline); support triage to specialist resolvers depends on the ticket (Q2 → supervisor); a report across legal, finance, and technical teams is multi-domain (Q3 → hierarchical); reviewing a contract clause from several angles needs debate (Q4 → moderated peer-to-peer). The supervisor route is just a fast-model classifier feeding dict.get(label, default) — keep that default mandatory so an unexpected label never crashes the system.

TIP
Prototype flat, then scale into shape
Start almost every project as a flat supervisor — easiest to observe, cheapest to change. Refactor toward a pipeline once the route stops varying, or a hierarchy once context gets crowded. Choose P2P for capability, never control: if you cannot name the debate the agents must have, you do not need it.

Questions & Answers

Q: My supervisor is a latency bottleneck on every request. Must I abandon the pattern?
No — separate routing from reasoning. Use a small, fast model for the supervisor (its job is just to emit a structured next-action), reserve stronger models for workers, and cache repetitive routing by input hash. Move to a hierarchy only when context size, not latency, is the constraint — a hierarchy adds hops and makes latency worse.
Q: How do I keep a P2P group chat from looping forever and draining my budget?
Three guardrails, all mandatory: a hard max_turns cap enforced in code (never trusted to the model), an explicit completion token (like TASK_COMPLETE) any agent can emit, and a per-conversation token budget that aborts when exceeded. A moderated chat — one selector picking the next speaker — also sharply reduces loop risk.
Q: A pipeline stage fails on item 8,000 of 10,000. How should partial failure behave?
Decide per item, not per batch. Each item carries an idempotency key so retrying item 8,000 does not reprocess the 7,999 that succeeded; route hard-failed items to a dead letter queue rather than aborting the run. This is why pipelines map so well onto workflow engines — Temporal and Prefect give per-activity retries and durable state for free (Lesson 6), and the recovery patterns built on top, like dead letter queues and circuit breakers, are Lesson 7.
Q: How does pattern choice affect what I can observe in production?
Heavily. Centralized patterns (supervisor, pipeline) have one place that knows the full state, so a single trace ID per request suffices. Peer-to-peer needs distributed tracing with span propagation across every agent-to-agent message, or you cannot reconstruct what happened. If observability is a hard requirement, prefer a centralized topology. Lesson 8 covers the instrumentation.

Key Takeaways

  1. Four topologies, one spectrum. Pipeline, supervisor, hierarchical, and peer-to-peer span centralized-deterministic to decentralized-adaptive.
  2. Start flat with a supervisor. It centralizes state, retries, and logging — cheapest to observe and change. Refactor only under concrete pressure.
  3. Determinism and adaptivity trade off. Runtime routing freedom is testability you give up. Buy it only where the workload demands.
  4. Cost follows context re-sends. P2P replays the transcript each turn (expensive); pipelines pass minimal payloads (cheap). Topology is your first cost lever.
  5. Patterns compose per boundary. Nest a supervisor inside a pipeline stage, a pipeline inside a hierarchy leaf — pick the simplest each boundary needs.
  6. Pipelines and workflow engines fit naturally. A staged DAG maps onto Temporal, Prefect, or Airflow, inheriting durable retries, timeouts, and DLQs.

Next Steps: Lesson 3: Communication Between Agents