Orchestration Patterns
Learning Outcomes
- Distinguish the four dominant orchestration patterns and their control-flow topologies
- Map each pattern's strengths and failure modes onto distributed-systems concerns you know
- Implement a minimal supervisor and a minimal pipeline in Python
- Evaluate latency, cost, and fault-tolerance trade-offs before committing to a topology
- Select the right pattern for a workload using a repeatable decision procedure
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Topology as the first architectural decision |
| Supervisor | 8 min | Central coordinator delegating to workers |
| P2P & hierarchical | 10 min | Mesh collaboration and recursive supervisors |
| Pipeline | 7 min | Sequential staged processing |
| Trade-offs | 6 min | Latency, cost, and fault tolerance compared |
| Selection | 4 min | A decision procedure with worked examples |
| Wrap-up | 2 min | Key takeaways, preview communication |
Before You Begin
Pre-work:
- Complete Lesson 1: Why Orchestration? — we assume you accept the motivation for multi-agent systems.
- Review tool use and agent loops from the Agentic AI course, particularly Agent Architectures. We treat each agent as a black-box function here.
Shopping List:
- Python 3.11+ with
pip install anthropicand anANTHROPIC_API_KEYexported - A scratch project directory and your editor of choice
A supervisor (orchestrator, router) is a single coordinating agent that owns the task graph: it takes the goal, decides which specialist acts next, dispatches a sub-task, collects the result, and repeats. Workers never talk to each other — every edge passes through the supervisor, a star topology with stateless workers at the points.
This is the workhorse pattern and where most teams should start: centralized routing makes the supervisor the single source of truth, so the system is easy to reason about, log, and debug. A minimal supervisor, where worker(system, task, model) is a one-line wrapper around client.messages.create(...) returning the text:
FAST, SMART = "claude-haiku-latest", "claude-sonnet-latest"
WORKERS = { # stateless specialists keyed by role
"search": lambda t: worker("Find relevant facts. Be terse.", t, FAST),
"writer": lambda t: worker("Draft clear prose.", t, SMART),
"reviewer": lambda t: worker("Critique drafts for accuracy.", t, SMART)}
def supervisor(goal: str) -> str: # route hard-coded for clarity
draft = WORKERS["writer"](goal + " / " + WORKERS["search"](goal))
return WORKERS["writer"]("Revise per: " + WORKERS["reviewer"](draft))
In production that route is itself a model call returning a structured action naming the next worker — made dynamic in Lesson 4: Task Decomposition at Scale.
In peer-to-peer, agents communicate directly with no central router; each decides who to address next from the conversation so far. This is the model behind frameworks like AutoGen — a "group chat" of specialists contributing until a termination condition fires.
[analyst] ◀──▶ [critic] mesh: any agent may
▲ ▲ address any other
└──[coder]────┘
Emergent collaboration is the strength — agents debate and correct each other, reaching solutions a fixed pipeline never would. The cost is a non-deterministic topology that can loop, so you must bound it with a hard turn cap and an explicit TASK_COMPLETE token. The loop iterates range(max_turns), picks a speaker (round-robin or an LLM moderator), and appends each reply to a shared transcript. Without that cap and token, two agents ping-pong forever — treat max_turns as an SLO guardrail, not a tuning knob.
Hierarchical orchestration is the supervisor pattern applied recursively: a top-level supervisor delegates to mid-level supervisors, each owning a team of workers. It is the org-chart of agent systems — keeping any one coordinator's span of control, here its context window, manageable.
[ EXECUTIVE ] tree topology
│ │
[RESEARCH lead] [WRITING lead] team leads
│ │ │ │
search summarize draft edit workers
The advantage over a flat supervisor is context isolation: the executive holds only the goal plus per-team summaries, each lead only its objective plus its workers' results, each worker only a single task. Every prompt stays small even when the workflow is large. The trade-off is depth-proportional latency — a leaf result bubbles all the way up, each layer adding a round-trip, so a three-layer hierarchy means roughly six sequential LLM calls before the executive responds.
A pipeline chains agents into fixed stages (extract ─▶ validate ─▶ transform ─▶ summarize), the output of stage N feeding stage N+1. No router makes dynamic decisions; the DAG is baked in at design time, the multi-agent analogue of a Unix pipe or ETL job.
Pipelines win when the work has a stable shape — document processing, data enrichment, content generation with QA. The runner is trivial: fold the payload through a list of named stages, wrapping each so a failure reports which stage broke.
def pipeline(stages: list[tuple[str, callable]], payload: str) -> str:
for name, run in stages:
try:
payload = run(payload)
except Exception as e:
raise RuntimeError("stage '" + name + "' failed: " + str(e)) from e
return payload
Fixed stage positions make pipelines the easiest pattern to test, cache, and parallelize. The DAG also maps onto a workflow engine — each stage a durable activity with its own retry and timeout, a Prefect @task with retries=3 giving backoff for free, the subject of Lesson 6: Workflow Engines.
The four patterns sit on a spectrum from centralized-deterministic (pipeline) to decentralized-emergent (peer-to-peer). Read this as a set of dials, not a leaderboard.
| Dimension | Supervisor | Peer-to-peer | Hierarchical | Pipeline |
|---|---|---|---|---|
| Control flow | Star | Mesh | Tree | Linear DAG |
| Latency | 2 hops/task | Unbounded* | Depth × hops | Sum of stages |
| SPOF | Supervisor | None | Subtree roots | Any stage |
| Observability | Easy | Hard | Medium | Easy |
| Token cost | Medium | High | Med-high | Predictable |
| Best for | Dynamic routing | Open-ended debate | Large decomposable work | Stable repeatable flows |
* Bounded only by your hard max_turns guardrail.
Two themes recur: determinism and adaptivity trade off, and centralization concentrates failure but simplifies recovery — a supervisor is a SPOF, but the single home for a circuit breaker and trace (Lesson 7), while cost is itself a topology decision (Lesson 9: Cost Management & Optimization).
For a new workload, walk these questions in order; the first "yes" is usually your answer.
Q1 Steps fixed and known in advance? ─▶ PIPELINE
Q2 Routing depends on runtime results? ─▶ SUPERVISOR
Q3 Supervisor prompt/tools too large, or sub-domains
needing distinct policies? ─▶ HIERARCHICAL
Q4 Task truly requires agents to debate / critique? ─▶ PEER-TO-PEER (moderated)
Applied to real workloads: nightly invoice extract + validate + post is fixed steps (Q1 → pipeline); support triage to specialist resolvers depends on the ticket (Q2 → supervisor); a report across legal, finance, and technical teams is multi-domain (Q3 → hierarchical); reviewing a contract clause from several angles needs debate (Q4 → moderated peer-to-peer). The supervisor route is just a fast-model classifier feeding dict.get(label, default) — keep that default mandatory so an unexpected label never crashes the system.
Questions & Answers
max_turns cap enforced in code (never trusted to the model), an explicit completion token (like TASK_COMPLETE) any agent can emit, and a per-conversation token budget that aborts when exceeded. A moderated chat — one selector picking the next speaker — also sharply reduces loop risk.Key Takeaways
- Four topologies, one spectrum. Pipeline, supervisor, hierarchical, and peer-to-peer span centralized-deterministic to decentralized-adaptive.
- Start flat with a supervisor. It centralizes state, retries, and logging — cheapest to observe and change. Refactor only under concrete pressure.
- Determinism and adaptivity trade off. Runtime routing freedom is testability you give up. Buy it only where the workload demands.
- Cost follows context re-sends. P2P replays the transcript each turn (expensive); pipelines pass minimal payloads (cheap). Topology is your first cost lever.
- Patterns compose per boundary. Nest a supervisor inside a pipeline stage, a pipeline inside a hierarchy leaf — pick the simplest each boundary needs.
- Pipelines and workflow engines fit naturally. A staged DAG maps onto Temporal, Prefect, or Airflow, inheriting durable retries, timeouts, and DLQs.
Next Steps: Lesson 3: Communication Between Agents