Orchestration Principles

Reference intermediate

A condensed set of operating principles for multi-agent systems. Each rule names a failure mode and the design move that mitigates it. Treat agents as unreliable distributed workers: non-deterministic, occasionally slow, capable of confidently wrong output. The discipline that makes microservices survivable — explicit contracts, bounded blast radius, observability, least privilege — applies directly.

Prefer the Simplest Pattern That Works

Most "multi-agent" problems are a sequential pipeline or a single supervisor with a few workers. Reserve peer-to-peer and hierarchical graphs for cases where you can articulate why a simpler topology fails.

Pattern When it is the right call Cost of choosing it too early
Single agent + tools Task fits one context window, one skill None — this is the baseline
Pipeline (chain) Fixed stages, clear data flow Rigid; hard to branch
Supervisor + workers Dynamic delegation, specialist skills One coordinator becomes a bottleneck
Peer-to-peer / graph Emergent collaboration, negotiation Loops, non-termination, no clear owner
# Start here. Add agents only when a single one demonstrably can't cope.
result = await agent.run(task)
if result.needs_specialist:
    result = await supervisor.delegate(task, workers=[search, synth])

See /courses/05-agent-orchestration/lesson-02/ for the full pattern catalogue.

Make Handoffs Explicit

A handoff is an API boundary; define a schema for what crosses it. Never pass an entire conversation transcript when a structured summary will do — token cost compounds and downstream agents over-anchor on irrelevant detail.

Handoff style Pass this Avoid
Structured Validated objects (Pydantic, JSON Schema) Free-text dumps
Summarised Distilled findings + provenance Raw tool output
Referenced Pointer/key to shared store Inlining large blobs
from pydantic import BaseModel

class ResearchHandoff(BaseModel):
    query: str
    findings: list[str]
    sources: list[str]       # provenance travels with the claim
    confidence: float        # let the next agent reason about trust

# Validate at the boundary, before the next agent ever sees it.
payload = ResearchHandoff.model_validate(coordinator_output)

Treat shared-state writes the same way: a blackboard is eventually consistent, so give every entry a key, a writer ID, and a version. See /courses/05-agent-orchestration/lesson-03/.

Design for Partial Failure

Some sub-tasks will succeed while others fail; the system must make progress with partial results and never leave a workflow wedged. Make agent steps idempotent so retries are safe, and decide explicitly what a partial result is worth.

Failure mode Mechanism Notes
Transient (timeout, 429) Retry with exponential backoff + jitter Cap attempts; respect a retry budget
Persistent (bad input) Dead-letter queue + escalate Don't retry forever
Degraded dependency Fallback agent / simpler model Return partial result with a flag
Repeated transient Circuit breaker Stop hammering a failing dependency
import random, asyncio

async def retry(call, attempts=4, base=0.5, budget_left=lambda: True):
    for i in range(attempts):
        if not budget_left():
            raise BudgetExhausted()
        try:
            return await call()                      # call must be idempotent
        except TransientError:
            if i == attempts - 1:
                raise
            await asyncio.sleep(base * 2**i + random.random())

Workflow engines (Temporal, Prefect, Airflow) give you durable retries, timeouts, and resumable state — wrap each non-deterministic agent call as an activity so a crash resumes from the last completed step rather than re-running the whole graph. See /courses/05-agent-orchestration/lesson-06/.

Bound Total Cost

Non-deterministic agents loop, re-plan, and re-fetch. Without hard limits a single run can consume unbounded tokens and dollars. Bound the dimensions below independently and fail closed when any is hit.

Dimension Cap Enforced where
Tokens Per-agent and per-run budget Orchestrator, before each call
Steps Max iterations / delegations Loop guard in the supervisor
Wall-clock Timeout per activity and per run Workflow engine
Money Spend ceiling across the run Cost meter, checked between steps
class Budget:
    def __init__(self, max_tokens, max_steps):
        self.tokens_left = max_tokens
        self.steps_left = max_steps

    def charge(self, used_tokens):
        self.tokens_left -= used_tokens
        self.steps_left -= 1
        if self.tokens_left <= 0 or self.steps_left <= 0:
            raise BudgetExhausted()   # fail closed, never loop on empty

Optimise within the budget by routing cheap tasks to faster/smaller models, caching repeated context (prompt caching), and caching identical sub-task results. See /courses/05-agent-orchestration/lesson-09/.

Keep Traces

A multi-agent run is a distributed trace; without correlation you cannot answer "which agent caused this output?" Propagate one run ID through every call, tool invocation, and handoff, and emit structured events — not prose logs.

Capture Why it matters
Run ID / trace ID on every span Correlate work across agents
Parent/child span IDs Reconstruct the delegation tree
Inputs, outputs, tool calls Replay and root-cause
Token usage + latency per span Cost and SLO attribution
Decision + reasoning summary Explain non-deterministic choices
from opentelemetry import trace
tracer = trace.get_tracer("orchestrator")

with tracer.start_as_current_span("agent.synthesis") as span:
    span.set_attribute("run.id", run_id)
    span.set_attribute("agent.name", "synthesis")
    out = await synth.run(payload)
    span.set_attribute("tokens.total", out.usage.total_tokens)

Persist enough to replay a failed run deterministically against recorded inputs. See /courses/05-agent-orchestration/lesson-08/.

Isolate Agent Permissions

Each agent should hold the narrowest set of tools and credentials its job requires; a compromised or hallucinating agent must not reach beyond its lane. Authenticate agent-to-agent calls and log every privileged action.

Control Rule
Tool scope Grant only the tools an agent needs; deny by default
Credentials Per-agent service identity, never a shared god-key
Data access Read/write only the partitions in scope
Side effects Gate writes, payments, and deletes behind explicit approval
Audit Log actor, action, target for every privileged call
search_agent = Agent(
    name="search",
    tools=[web_search],          # read-only, no filesystem, no payments
    instructions="Retrieve and cite. Never write to shared stores.",
)
writer_agent = Agent(
    name="writer",
    tools=[write_doc],           # scoped to one output location
    instructions="Persist the approved draft only.",
)

Least-privilege tooling, agent-to-agent authentication, and audit logging are production gates covered in /courses/05-agent-orchestration/lesson-10/.

Related References

Resource Path
Course landing /courses/05-agent-orchestration/
Prerequisite (single agent) /courses/04-agentic-ai/