Why Orchestration?

30 min beginner Lesson 1

Learning Outcomes

  • Identify the four structural limits of single-agent systems: context, complexity, specialization, and reliability
  • Diagnose when a task has outgrown a single agent and needs decomposition
  • Compare the failure modes of monolithic agents against coordinated multi-agent systems
  • Map a real workflow into discrete, specialized agent roles with clear handoffs
  • Evaluate the cost, latency, and reliability trade-offs that orchestration introduces

Lesson Plan

Segment Duration Topic
Intro 3 min Where single agents stop scaling
Limit 1 5 min Context window as a shared budget
Limit 2 5 min Complexity ceilings and reasoning drift
Limit 3 5 min Specialization vs. the generalist prompt
Limit 4 4 min Reliability and compounding error rates
Case study 5 min A document pipeline, decomposed
Trade-offs 2 min What orchestration costs you
Wrap-up 1 min Key takeaways, preview next lesson

Before You Begin

Pre-work:

  • Complete the Agentic AI course, or be comfortable building a single tool-using agent. This lesson assumes you know agent loops, tool use (Agentic AI Lesson 3), and memory (Agentic AI Lesson 4).
  • Note where your own agent's system prompt has grown past a screen — that growth is the symptom.

Shopping List:

  • Python 3.11+, an Anthropic API key, and the SDK (pip install anthropic)
  • A real, slightly-too-big task you have tried with one agent (a triage bot, research assistant, or code reviewer)

1 The Context Window Is a Shared Budget, Not a Cache

A single agent crams everything into one context window — system prompt, tools, history, documents, reasoning — all competing for one fixed budget. As a task runs longer you hit two walls: the hard one (you run out of tokens) and the insidious one, degradation well before the limit, where the model attends less reliably to information buried mid-context. The cost side shows up the moment you instrument it:

resp = client.messages.create(
    model="claude-sonnet-4-6", max_tokens=1024,
    messages=messages,            # full history, re-sent every turn
)
print(resp.usage.input_tokens)    # climbs monotonically

Because every turn re-sends the full history, input_tokens climbs monotonically — a 30-step task pays roughly the sum of an arithmetic series, not 30 flat turns, the quiet cost driver behind most "why is my agent so expensive?" tickets. Orchestration instead partitions the budget: each worker gets a fresh window scoped to its sub-task — never another worker's raw logs, only a structured handoff.

WARNING
Watch Out
Do not orchestrate the moment context feels tight. Try compaction and retrieval first; split only when concerns are truly separable.

2 Complexity Ceilings and Reasoning Drift

A single agent asked to plan, execute, verify, and report does all four poorly at the margin: it juggles competing objectives in one pass, and its attention to the goal drifts as the trajectory lengthens. The tell is a system prompt that has become a spec document — branching rules, "if X then do Y unless Z." A prompt that reads like a state machine asks one inference to be one, and it will not hold the invariant.

Decomposition moves the coordination logic out of the model into deterministic code — testable and drift-free:

WORKERS = {"research": research_agent, "compute": compute_agent}

def orchestrate(goal: str):
    plan = planner_agent(goal)                 # one focused agent
    results = [WORKERS[t.kind](t) for t in plan.subtasks]
    return synthesis_agent(goal, results)      # another focused agent

The control flow is ordinary Python you can unit-test, and each agent does one thing well. This is the heart of the argument: push determinism into code, keep judgment in the agents (Lesson 2).

TIP
Heuristic
If you cannot describe an agent's job in one sentence without the word and, it is two agents.

3 Specialization Beats the Generalist Prompt

A generalist carries every tool, instruction, and edge case at once; a specialist carries only what its role needs — a sharper prompt, a narrower tool surface, and the right model per role. Extraction is cheap, fast, high-volume; deep synthesis is rare and expensive. A single agent forces one model for both — orchestration lets you route by difficulty.

from dataclasses import dataclass

@dataclass
class AgentSpec:
    model: str          # match capability to task difficulty
    system: str         # focused, single-responsibility prompt
    tools: list[dict]   # least-privilege tool surface

# Cheap, high-volume, no tools — pure transformation:
extractor = AgentSpec("claude-haiku-4-5",
    "Extract invoice fields as strict JSON. Do not infer.", tools=[])
# Capable, low-volume, one read-only tool:
reviewer = AgentSpec("claude-opus-4-8",
    "Audit fields against the source for hallucinations.",
    tools=[{"name": "fetch_source_doc"}])

The extractor has no tools — it cannot reach the network or filesystem, so the blast radius of a bad output is near zero. The reviewer gets one read-only tool. This is least-privilege design at the agent boundary, the same principle as a microservice's IAM role.

NOTE
Key Insight
Specialization is about safety and economics, not just quality: narrow tools shrink what can go wrong, and per-role model choice is the biggest cost lever.

4 Reliability: Why Error Rates Compound

This is the limit that bites in production. Suppose each step succeeds 95% of the time — generous for an LLM. A single agent chaining ten such steps with no checkpoints has an end-to-end success of:

0.95 ** 10 = 0.5987...   # roughly a 60% chance of finishing cleanly

Four in ten runs fail, and as one long trajectory you often cannot tell which step failed or resume. Multi-agent systems do not raise per-step reliability; they give you checkpointing and isolation. Each handoff is a durable boundary (keyed by sub-task ID for idempotency) where you validate, retry, fall back, or escalate without discarding what succeeded — a validate-then-act gate:

def safe_handoff(produce, validate, fallback, attempts=3):
    for _ in range(attempts):
        out = produce()
        if validate(out):
            return out
    return fallback()   # degrade gracefully instead of crashing

The same fault tolerance you already apply to distributed systems — bulkheads, retries, dead-letter queues — now applied to AI steps. Lesson 7 goes deep.

TIP
Rule of thumb
Multiply per-step success rates to estimate end-to-end reliability. Below your SLO, you need checkpoints — which means boundaries — which means orchestration.

5 Case Study — A Document Processing Pipeline

A finance team processes invoices: extract line items, validate them against a purchase-order database, flag anomalies, and summarize for an approver. As one agent, the prompt balloons to cover all four jobs and needs database, PDF, and write tools at once — and a failure on invoice 7 of 200 kills the batch. Decomposed, a coordinator fans out to four specialists, each with a least-privilege tool surface:

Agent Model tier Tools Handoff out
Extractor Cheap None Structured JSON fields
Validator Cheap Read-only PO lookup Validated record + diffs
Anomaly checker Capable Read-only history Flags with reasons
Summarizer Mid None Approver paragraph

The coordinator runs invoices independently, so a failure on invoice 7 retries only invoice 7, and you can swap the extractor's model without touching the summarizer.

NOTE
Key Insight
The win is not that four agents are smarter than one. Decomposition trades intelligence for operability: each bounded agent is tested, retried, traced, and priced on its own.

This shape recurs across domains; we build it with the Anthropic Agent SDK in Lesson 5.


6 The Trade-offs — Orchestration Is Not Free

Orchestration is a distributed system and inherits every hard problem of one. Latency: every handoff is a round trip, so five sequential agents mean five sequential model calls. Consistency: two agents on a shared workspace can act on stale state, like services reading a stale replica (Lesson 3). A simple gate before you orchestrate:

def should_orchestrate(task) -> bool:
    return (
        task.separable_concerns >= 2     # distinct responsibilities
        and task.estimated_steps >= 6    # checkpoints start to matter
        and task.needs_per_step_recovery # partial failure must survive
    )   # all false? a single well-instrumented agent wins

Orchestration earns its complexity only against a limit from steps 1–4 that is blocking you today.

WARNING
Watch Out
The most expensive multi-agent systems are built before a single agent was proven insufficient. Orchestrate against a measured limit, not a guess.

Questions & Answers

Q: Bigger context windows keep shipping. Won't that kill the context argument?
A bigger window raises the hard limit, not the soft one — effective attention still degrades across very long contexts. And it does nothing for the other three limits: drift, specialization, and compounding error rates, all independent of window size.
Q: Isn't multi-agent just strictly more expensive — more calls, more tokens?
Not necessarily. A single agent re-bills its entire growing transcript every turn, which dominates cost on long tasks. Partitioning into small per-agent windows often *lowers* total input tokens, and routing easy steps to cheaper models lowers it further. It costs more only if you duplicate context across agents — hence handoff design and cost management (Lessons 3 and 9).
Q: How do I debug when something goes wrong across agents?
Assign a correlation ID at the coordinator and propagate it through every handoff, then emit a span per agent call — the same distributed-tracing discipline you use for microservices. Each handoff is a natural place to log input and output, so failures localize to one agent rather than an opaque transcript (Lesson 8).
Q: Can agents get stuck talking to each other in a loop?
Yes — a real hazard in peer-to-peer topologies. Guard against it with hard limits: a max hop count per request, a turn budget, and idempotency keys so a re-delivered message does not trigger fresh work. Supervisor and pipeline patterns avoid most loops by construction, since control flow lives in code, not agent chatter (Lessons 2 and 7).

Key Takeaways

  1. Four structural limits drive orchestration: shared context budget, complexity drift, specialization needs, and compounding error rates. If none blocks you, you do not need multiple agents yet.
  2. Context is a budget to partition, not a cache to grow: a single agent re-bills its whole transcript each turn; orchestration gives each worker a private window.
  3. Push determinism into code, keep judgment in agents: logic that reads like a state machine belongs in testable control flow, not one inference pass.
  4. Boundaries enable reliability: per-step success rates multiply, so long chains fail often — handoffs add the checkpoints and isolation that make a workflow recoverable.
  5. Orchestration is a distributed system: it inherits latency, eventual consistency, and loop hazards. Establish a single-agent baseline and an SLO first, then orchestrate against a measured limit.

Next Steps: Lesson 2: Orchestration Patterns