Orchestration Principles
A condensed set of operating principles for multi-agent systems. Each rule names a failure mode and the design move that mitigates it. Treat agents as unreliable distributed workers: non-deterministic, occasionally slow, capable of confidently wrong output. The discipline that makes microservices survivable — explicit contracts, bounded blast radius, observability, least privilege — applies directly.
Prefer the Simplest Pattern That Works
Most "multi-agent" problems are a sequential pipeline or a single supervisor with a few workers. Reserve peer-to-peer and hierarchical graphs for cases where you can articulate why a simpler topology fails.
| Pattern | When it is the right call | Cost of choosing it too early |
|---|---|---|
| Single agent + tools | Task fits one context window, one skill | None — this is the baseline |
| Pipeline (chain) | Fixed stages, clear data flow | Rigid; hard to branch |
| Supervisor + workers | Dynamic delegation, specialist skills | One coordinator becomes a bottleneck |
| Peer-to-peer / graph | Emergent collaboration, negotiation | Loops, non-termination, no clear owner |
# Start here. Add agents only when a single one demonstrably can't cope.
result = await agent.run(task)
if result.needs_specialist:
result = await supervisor.delegate(task, workers=[search, synth])
See /courses/05-agent-orchestration/lesson-02/ for the full pattern catalogue.
Make Handoffs Explicit
A handoff is an API boundary; define a schema for what crosses it. Never pass an entire conversation transcript when a structured summary will do — token cost compounds and downstream agents over-anchor on irrelevant detail.
| Handoff style | Pass this | Avoid |
|---|---|---|
| Structured | Validated objects (Pydantic, JSON Schema) | Free-text dumps |
| Summarised | Distilled findings + provenance | Raw tool output |
| Referenced | Pointer/key to shared store | Inlining large blobs |
from pydantic import BaseModel
class ResearchHandoff(BaseModel):
query: str
findings: list[str]
sources: list[str] # provenance travels with the claim
confidence: float # let the next agent reason about trust
# Validate at the boundary, before the next agent ever sees it.
payload = ResearchHandoff.model_validate(coordinator_output)
Treat shared-state writes the same way: a blackboard is eventually consistent, so give every entry a key, a writer ID, and a version. See /courses/05-agent-orchestration/lesson-03/.
Design for Partial Failure
Some sub-tasks will succeed while others fail; the system must make progress with partial results and never leave a workflow wedged. Make agent steps idempotent so retries are safe, and decide explicitly what a partial result is worth.
| Failure mode | Mechanism | Notes |
|---|---|---|
| Transient (timeout, 429) | Retry with exponential backoff + jitter | Cap attempts; respect a retry budget |
| Persistent (bad input) | Dead-letter queue + escalate | Don't retry forever |
| Degraded dependency | Fallback agent / simpler model | Return partial result with a flag |
| Repeated transient | Circuit breaker | Stop hammering a failing dependency |
import random, asyncio
async def retry(call, attempts=4, base=0.5, budget_left=lambda: True):
for i in range(attempts):
if not budget_left():
raise BudgetExhausted()
try:
return await call() # call must be idempotent
except TransientError:
if i == attempts - 1:
raise
await asyncio.sleep(base * 2**i + random.random())
Workflow engines (Temporal, Prefect, Airflow) give you durable retries, timeouts, and resumable state — wrap each non-deterministic agent call as an activity so a crash resumes from the last completed step rather than re-running the whole graph. See /courses/05-agent-orchestration/lesson-06/.
Bound Total Cost
Non-deterministic agents loop, re-plan, and re-fetch. Without hard limits a single run can consume unbounded tokens and dollars. Bound the dimensions below independently and fail closed when any is hit.
| Dimension | Cap | Enforced where |
|---|---|---|
| Tokens | Per-agent and per-run budget | Orchestrator, before each call |
| Steps | Max iterations / delegations | Loop guard in the supervisor |
| Wall-clock | Timeout per activity and per run | Workflow engine |
| Money | Spend ceiling across the run | Cost meter, checked between steps |
class Budget:
def __init__(self, max_tokens, max_steps):
self.tokens_left = max_tokens
self.steps_left = max_steps
def charge(self, used_tokens):
self.tokens_left -= used_tokens
self.steps_left -= 1
if self.tokens_left <= 0 or self.steps_left <= 0:
raise BudgetExhausted() # fail closed, never loop on empty
Optimise within the budget by routing cheap tasks to faster/smaller models, caching repeated context (prompt caching), and caching identical sub-task results. See /courses/05-agent-orchestration/lesson-09/.
Keep Traces
A multi-agent run is a distributed trace; without correlation you cannot answer "which agent caused this output?" Propagate one run ID through every call, tool invocation, and handoff, and emit structured events — not prose logs.
| Capture | Why it matters |
|---|---|
| Run ID / trace ID on every span | Correlate work across agents |
| Parent/child span IDs | Reconstruct the delegation tree |
| Inputs, outputs, tool calls | Replay and root-cause |
| Token usage + latency per span | Cost and SLO attribution |
| Decision + reasoning summary | Explain non-deterministic choices |
from opentelemetry import trace
tracer = trace.get_tracer("orchestrator")
with tracer.start_as_current_span("agent.synthesis") as span:
span.set_attribute("run.id", run_id)
span.set_attribute("agent.name", "synthesis")
out = await synth.run(payload)
span.set_attribute("tokens.total", out.usage.total_tokens)
Persist enough to replay a failed run deterministically against recorded inputs. See /courses/05-agent-orchestration/lesson-08/.
Isolate Agent Permissions
Each agent should hold the narrowest set of tools and credentials its job requires; a compromised or hallucinating agent must not reach beyond its lane. Authenticate agent-to-agent calls and log every privileged action.
| Control | Rule |
|---|---|
| Tool scope | Grant only the tools an agent needs; deny by default |
| Credentials | Per-agent service identity, never a shared god-key |
| Data access | Read/write only the partitions in scope |
| Side effects | Gate writes, payments, and deletes behind explicit approval |
| Audit | Log actor, action, target for every privileged call |
search_agent = Agent(
name="search",
tools=[web_search], # read-only, no filesystem, no payments
instructions="Retrieve and cite. Never write to shared stores.",
)
writer_agent = Agent(
name="writer",
tools=[write_doc], # scoped to one output location
instructions="Persist the approved draft only.",
)
Least-privilege tooling, agent-to-agent authentication, and audit logging are production gates covered in /courses/05-agent-orchestration/lesson-10/.
Related References
| Resource | Path |
|---|---|
| Course landing | /courses/05-agent-orchestration/ |
| Prerequisite (single agent) | /courses/04-agentic-ai/ |