Production Multi-Agent Systems

60 min advanced Lesson 10

Learning Outcomes

  • Design a deployment topology that scales orchestrators and worker agents independently
  • Define SLOs and error budgets that account for non-deterministic agent behaviour
  • Run chaos experiments and implement graceful degradation paths for partial failures
  • Control agent-to-agent access with authentication, least-privilege tools, and audit logging
  • Produce a production-readiness checklist and incident-response runbook for a multi-agent system

Lesson Plan

Segment Duration Topic
Intro 3 min From prototype to production — what changes
Architecture 9 min Deployment topology and independent scaling
Reliability 10 min SLOs, error budgets, and idempotency
Chaos 9 min Fault injection and graceful degradation
Access control 10 min Agent-to-agent auth, least privilege, audit logs
Operations 8 min Deployment strategy and rollout safety
Runbooks 8 min Incident response and the readiness checklist
Wrap-up 3 min Key takeaways, preview Local LLMs

Before You Begin

Pre-work:

Shopping List:

  • A container runtime (Docker) and a target — Kubernetes, ECS, or a managed runner
  • A workflow engine (Temporal, Prefect, or Airflow) from Lesson 6
  • An OpenTelemetry collector and a metrics backend (from Lesson 8)
  • A secrets manager (Vault, AWS Secrets Manager, or cloud equivalent)
  • Python 3.11+ with the Anthropic SDK installed

1 From Prototype to Production — What Actually Changes

Your prototype likely ran the orchestrator and every agent in one process, on one machine, with a single API key and only the retries you added by hand. Production is a different system: it must survive a worker that hangs mid-generation, an endpoint that returns a 529, a deploy that ships a bad prompt, and a 3am page with no human who understands the agent graph.

The shift is from "does it work once" to "does it keep working under load, partial failure, and change." Four properties matter more than feature count:

Property Prototype reality Production requirement
Availability Works when you run it SLO with an error budget
Recoverability Restart and pray Durable state, replayable runs
Isolation One process, one key Independent scaling, scoped credentials
Observability print() and logs Distributed traces and alerts

The single most important architectural decision is to separate durable orchestration state from agent compute. The orchestrator must never hold the only copy of "where are we in this run." That belongs in a workflow engine or database so any worker can crash without losing the run. Agents are just unusually expensive, slow, non-deterministic workers — the distributed-systems rules still apply.

NOTE
Key Insight
Treat each agent as a stateless worker and the orchestrator as a thin scheduler over durable state. If a worker dying loses progress, you have a prototype, not a production system.

2 Deployment Topology — Scaling Agents Independently

Different agents have different cost and latency profiles. A coordinator does cheap routing; a synthesis agent burns large context windows; a fact-checker fans out into many small calls. Share one deployment and you scale the cheap thing to feed the expensive thing. Deploy each agent role as its own service so you can size and scale them independently. A queue between orchestrator and workers gives you back-pressure, turning the model API's rate limits into queue depth instead of a cascade of failures.

                 ┌──────────────┐
   request ─────▶│ Orchestrator │  (durable state in workflow engine)
                 └──────┬───────┘
                        │ enqueue tasks
              ┌─────────▼──────────┐
              │   task queue (SQS  │  ← back-pressure lives here
              │   / Redis / NATS)  │
              └───┬─────┬──────┬───┘
        ┌─────────▼─┐ ┌─▼──────┐ ┌▼────────────┐
        │ search    │ │ synth  │ │ fact-check  │  ← scaled separately
        │ agent x6  │ │ agent  │ │ agent x3    │
        │ (cheap)   │ │ x2     │ │             │
        └───────────┘ └────────┘ └─────────────┘

Build one container image and select the role at runtime via an AGENT_ROLE environment variable, so you ship a single artifact and configure topology at deploy time. On Kubernetes, give each role its own Deployment and autoscale on queue depth, not CPU — agent workers are I/O-bound on the model API, so CPU autoscaling lags badly.

# search-agent.deployment.yaml (excerpt)
spec:
  replicas: 6                              # tuned per role, autoscaled by KEDA
  template:
    spec:
      containers:
        - name: worker
          image: registry.example.com/agents:1.8.0
          env:
            - { name: AGENT_ROLE, value: "search" }   # one image, role via env
            - name: ANTHROPIC_API_KEY                  # per-role key, not shared
              valueFrom: { secretKeyRef: { name: search-creds, key: api-key } }
TIP
Scale on queue depth
Drive autoscaling from queue lag (messages waiting per worker), not CPU. A KEDA ScaledObject on the SQS or Redis queue length reacts in seconds; CPU-based HPA reacts after the latency damage is done.
WARNING
One key per role
Do not share a single API key across every agent. Per-role keys let you attribute spend, rate-limit a runaway agent in isolation, and revoke one role's access without taking the whole system down.

3 SLOs, Error Budgets, and the Non-Determinism Problem

You cannot promise "the agent always answers correctly" — the system is non-deterministic and the model is fallible. Define SLOs on measurable, externally meaningful properties, not on correctness you can't observe in real time. Good SLI candidates:

SLI Definition Target
Availability Runs completing without an unrecoverable error 99.5%
Latency End-to-end run time at p95 < 90s
Task success Runs passing automated validation (eval gates) 95%
Escalation rate Runs that fall back to a human < 5%
Cost per run p95 token spend per completed run in budget

Turn these into an error budget with burn-rate alerts: a fast burn (roughly 14x over 1h) pages on-call, a slow burn (6x over 6h) opens a ticket. The budget is permission to take risk — burning it fast freezes risky deploys; budget to spare means ship faster.

The reliability primitive that makes any of this safe is idempotency. Retries are guaranteed in a non-deterministic system, so every task carries an idempotency key and side-effecting tools dedupe on it — otherwise a retried "send email" agent sends the email twice.

def run_task(task):
    key = task["idempotency_key"]          # stable per logical task
    if (cached := results.get(key)):       # durable store, not in-memory
        return cached                      # replay-safe: never re-execute
    out = agent.invoke(task)
    results.put(key, out, ttl=DAYS_7)
    return out
NOTE
SLOs over accuracy
You operate on what you can measure in production. Correctness is validated offline with evals; SLOs govern availability, latency, success rate, and escalation. Conflating the two leads to alerts you can never act on.

4 Chaos Testing and Graceful Degradation

A multi-agent system has more failure modes than a single service: a worker hangs, the queue backs up, an endpoint degrades, a tool returns garbage, the context grows until a call is rejected. Find these in a controlled experiment, not in the incident. Chaos testing injects failures deliberately and verifies the system degrades instead of collapses. Inject faults at the seams you own — the tool layer and the queue:

# chaos.py — wrap tool calls with controlled fault injection
import random, time

class ChaosWrapper:
    def __init__(self, tool, fail_rate=0.0, latency_ms=0):
        self.tool, self.fail_rate, self.latency_ms = tool, fail_rate, latency_ms

    def __call__(self, *args, **kwargs):
        if random.random() < self.fail_rate:
            raise TimeoutError("chaos: injected tool failure")
        if self.latency_ms:
            time.sleep(self.latency_ms / 1000)
        return self.tool(*args, **kwargs)

# Staging only: search_tool = ChaosWrapper(search_tool, fail_rate=0.2, latency_ms=800)

Run each experiment with an explicit hypothesis and a blast-radius limit — staging first, then a small traffic slice. Common experiments:

Experiment Inject Expected graceful behaviour
Worker death Kill a pod mid-run Run resumes on another worker from durable state
Slow model +5s latency on one role Timeout fires, circuit breaker opens, fallback path used
Tool outage 100% failure on one tool Agent skips optional tool, marks result low-confidence
Queue flood 10x enqueue rate Back-pressure holds, autoscaler adds workers, no data loss

Graceful degradation means defining, ahead of time, what a partial answer looks like. If the fact-check agent is down, return the synthesis with a verified: false flag rather than failing the run. Build this on the circuit breaker and fallback patterns from Lesson 7.

async def orchestrate(task):
    draft = await synth_agent.run(task)
    try:
        checked = await with_timeout(factcheck_agent.run(draft), seconds=20)
        return {"answer": checked, "verified": True}
    except (TimeoutError, CircuitOpen):
        DEGRADED.inc()                     # metric -> alert if sustained
        return {"answer": draft, "verified": False, "degraded": True}
WARNING
Degradation must be visible
A silent fallback hides a failing dependency until it is total. Always emit a metric and tag the response when you degrade, so a rising degraded-response rate pages you before users notice.
TIP
Game day before launch
Run a scheduled chaos exercise (a 'game day') in staging before every major release. The goal is not to break things — it is to confirm your runbooks and fallbacks still match the system you actually built.

5 Access Control and Audit Between Agents

Once agents run as separate services, the messages between them cross a network. Treat agent-to-agent calls like any service-to-service traffic: authenticated, authorized, least-privilege. The principal hierarchy from the Agentic AI safety lesson still applies — but now there are many actors, and one confused agent can drive others.

Authentication. Every message carries a verifiable identity: mutual TLS between services plus short-lived signed tokens scoped to the calling agent's role. Verify the signature and role before acting on any task.

# auth.py — verify a signed agent-to-agent envelope
import jwt   # short-lived per-role tokens, rotated by the secrets manager

def verify_envelope(token: str, expected_audience: str) -> dict:
    claims = jwt.decode(
        token, PUBLIC_KEY, algorithms=["RS256"],
        audience=expected_audience,         # this worker's role
    )
    if claims["role"] not in ALLOWED_CALLERS[expected_audience]:
        raise PermissionError(f"{claims['role']} may not call {expected_audience}")
    return claims

Least-privilege tools. An agent holds only the tools its job requires: the search agent gets read-only web access, only the publishing agent writes to the CMS, no agent gets broad shell access. Encode this as a per-role allowlist enforced at the tool dispatcher — never by trusting the model to stay in bounds.

TOOL_GRANTS = {
    "search":     {"web_search", "read_docs"},
    "synth":      {"read_docs"},
    "factcheck":  {"web_search", "read_docs"},
    "publish":    {"cms_write"},            # only this role can write
}

def dispatch_tool(role: str, name: str, args: dict):
    if name not in TOOL_GRANTS[role]:
        AUDIT.log(role=role, tool=name, decision="DENIED", args=args)
        raise PermissionError(f"{role} not granted {name}")
    AUDIT.log(role=role, tool=name, decision="ALLOWED", args=args)
    return TOOLS[name](**args)

Audit logging. Log every tool call, handoff, and decision with run ID, agent role, and inputs — append-only, tamper-evident, correlated with the traces from Lesson 8. When something goes wrong (or a regulator asks), you must reconstruct exactly which agent did what and why.

WARNING
Tool output is untrusted input
Content an agent fetches — web pages, documents, other agents' messages — can contain injection attempts. Never let fetched text silently expand an agent's privileges. Keep instructions and data on separate channels and re-check authorization at the tool boundary.
NOTE
Defence in depth
Authentication stops impersonation; least privilege limits blast radius when a prompt injection succeeds anyway; the audit log lets you contain and explain it afterward. You need all three — no single control is sufficient.

6 Deployment Strategy and Safe Rollout

In an agent system, your "code" includes prompts, tool definitions, and model selection — all of which change behaviour without changing a line of Python. Version everything that affects output and roll it out behind the same safeguards you use for code: canaries, gradual ramp, automatic rollback on SLO regression. Pin prompts and models as deployable config, not hardcoded strings, so you can roll them back independently of the binary.

# config — versioned and shipped as artifacts
AGENT_CONFIG = {
    "search": {
        "prompt_version": "search-v7",      # stored in a prompt registry
        "model": "claude-sonnet-latest",    # name an alias, not a frozen id
        "max_tokens": 4096,
        "tool_grants": ["web_search", "read_docs"],
    },
}

Wrap the agent step in a workflow engine activity (the Temporal @activity.defn decorator, a Prefect task, or an Airflow operator) so retries, timeouts, and heartbeats are managed durably — a non-deterministic agent call becomes a replayable step that survives worker restarts, and the activity is the unit you version. The activity loads its config at call time, so a canary can serve a new prompt_version to a slice of traffic. Roll out a prompt or model change like a canary deploy:

Stage Traffic Gate to advance
Shadow 0% (compare only) Offline eval parity vs current
Canary 5% No SLO regression, cost within budget for 30 min
Ramp 25% → 50% Escalation and degraded rates stable
Full 100% Burn rate normal for 1h
TIP
Prompts are deploys
A prompt edit can change behaviour as much as a code change — and is far easier to ship carelessly. Put prompts under version control, run them through the same canary and eval gates, and keep the previous version one config flip away.
WARNING
Pin to aliases, freeze in incidents
Reference a model alias day-to-day so you get improvements, but be ready to pin an exact version during an incident. A silently upgraded model can shift behaviour under you — your eval gate should catch it before users do.

7 Runbooks and the Production-Readiness Checklist

The last mile is operational. When the system pages at 3am, the responder needs a runbook, not a research project. Write one per high-probability failure, each with a symptom, a diagnosis path, and concrete mitigation steps.

RUNBOOK: degraded-response rate > 10% for 5 min
  SYMPTOM   alert "agent.degraded.rate high"; users get verified:false answers
  DIAGNOSE  1. Check which role is degrading (dashboard: per-role error %)
            2. Check model provider status page + 5xx/529 rate
            3. Check circuit-breaker state for that role
  MITIGATE  a. If provider degraded: pin to fallback model alias
            b. If one role: scale that role, drain bad pods
            c. If queue backed up: confirm KEDA scaled; raise max replicas
  ROLLBACK  If correlated with last deploy: flip prompt_version to previous
  ESCALATE  If unrecovered in 15 min, page service owner; freeze deploys

Before you call the system production-ready, walk this checklist. It is the deliverable for this lesson — adapt it and keep it in the repo.

Category Check
Scaling Each role scales independently on queue depth; rate limits map to back-pressure
Reliability Durable run state survives any worker death; side-effecting tasks are idempotent; SLOs wired to burn-rate alerts
Resilience Circuit breakers and fallbacks for every dependency; degradation emits metrics; chaos game day passed
Access control mTLS plus signed role tokens; per-role tool allowlist at the dispatcher; append-only audit log; per-role secrets, rotated
Observability Distributed tracing across agents (L8); cost per run budgeted (L9)
Rollout Prompts and models versioned and canaried; automatic rollback on SLO regression
Ops A runbook per top failure mode; on-call rotation, escalation policy, and dashboards live
NOTE
Readiness is a gate, not a vibe
Run the checklist as an explicit go/no-go before launch and again before any major change. Every unchecked box is either a tracked risk you've accepted on the record or a blocker. There is no third option.
TIP
Write the runbook from the incident
After every real incident, update the runbook with what actually helped. Over a few cycles your runbooks converge on the real failure modes of your system instead of the ones you imagined.

Questions & Answers

Q: How do I set an availability SLO when the model itself is non-deterministic?
Separate availability from correctness. Your SLO covers what you control end-to-end — did the run complete without an unrecoverable error, within the latency target. Correctness is governed by an offline eval suite plus a "task success" SLI measured against automated validation and the escalation rate. You promise the system stays up and responsive, not that the model is always right — and you never write an SLO you cannot measure in production.
Q: Retries on agents scare me — won't I double-charge or send duplicate side effects?
Yes, unless every task is idempotent. Give each logical task a stable idempotency key, check a durable store before executing, and have side-effecting tools (emails, payments, writes) dedupe on that key. Read-only calls are cheap to retry; the danger is purely in the side effects, so guard those. A workflow engine like Temporal makes this the default by making activities replayable.
Q: Is per-agent containerization overkill? Can't I just run them in one process?
One process is fine until your agents have very different cost/latency profiles, or one misbehaving agent takes the whole system down. Independent deployment buys separate scaling, failure domains, credentials, and rate limits. If your agents are uniform and low-volume, a single deployment with internal concurrency is a reasonable start — just keep durable state external so you can split later without a rewrite.
Q: How do I stop text an agent reads from changing what it does?
Keep instructions and data apart, and limit what any one agent can do. Content an agent fetches — a page, a document, another agent's message — is material to work on, never a new instruction, so carry it on a separate channel from the system prompt and check permissions at the tool boundary rather than trusting what the text says. Then enforce a per-role tool allowlist at the dispatcher, so an agent that is led astray still cannot reach tools its role was never granted. Authentication stops impersonation; least privilege limits the damage when the first layer does not hold.
Q: How do I roll back a bad change when the "change" was just a prompt edit?
Treat prompts as versioned artifacts shipped as config, not strings baked into the image. Store them in a registry with a version id, reference that id from deployable config, and gate changes through the same canary stages as code. Rollback is then flipping prompt_version to the previous value — no rebuild, no redeploy. The same applies to model selection: reference an alias normally, but be able to pin an exact version instantly during an incident.

Key Takeaways

  1. Externalize durable state. The orchestrator is a thin scheduler over state held in a workflow engine or database. If a dying worker loses progress, you have a prototype, not a production system.
  2. Scale agent roles independently on queue depth. Different agents have different cost and latency profiles; give each its own deployment, its own key, and back-pressure via a queue so rate limits never cascade.
  3. Write SLOs you can measure, backed by error budgets. Govern availability, latency, success rate, escalation, and cost — not model correctness, which belongs to offline evals. Idempotency makes the inevitable retries safe.
  4. Engineer for partial failure. Chaos-test the seams, define what a degraded answer looks like before you need it, and make every fallback emit a metric so silent degradation can't hide a dying dependency.
  5. Authenticate and least-privilege every agent. mTLS plus signed role tokens, a per-role tool allowlist enforced at the dispatcher, fetched content treated as data rather than instructions, and an append-only audit log correlated with traces.
  6. Ship prompts and models like code. Version them, canary them, gate on SLOs, and keep rollback one config flip away — then run the readiness checklist as an explicit go/no-go before every launch.

Next Steps: Continue with Local LLMs