Production Multi-Agent Systems
Learning Outcomes
- Design a deployment topology that scales orchestrators and worker agents independently
- Define SLOs and error budgets that account for non-deterministic agent behaviour
- Run chaos experiments and implement graceful degradation paths for partial failures
- Control agent-to-agent access with authentication, least-privilege tools, and audit logging
- Produce a production-readiness checklist and incident-response runbook for a multi-agent system
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | From prototype to production — what changes |
| Architecture | 9 min | Deployment topology and independent scaling |
| Reliability | 10 min | SLOs, error budgets, and idempotency |
| Chaos | 9 min | Fault injection and graceful degradation |
| Access control | 10 min | Agent-to-agent auth, least privilege, audit logs |
| Operations | 8 min | Deployment strategy and rollout safety |
| Runbooks | 8 min | Incident response and the readiness checklist |
| Wrap-up | 3 min | Key takeaways, preview Local LLMs |
Before You Begin
Pre-work:
- Complete Lesson 7 — Error Handling & Recovery and Lesson 8 — Monitoring & Observability
- Review Lesson 9 — Cost Management & Optimization — production cost guardrails build on it
- Have a working multi-agent system from Lesson 5 or Lesson 6 you can deploy
- Recommended: revisit agent safety from the Agentic AI course
Shopping List:
- A container runtime (Docker) and a target — Kubernetes, ECS, or a managed runner
- A workflow engine (Temporal, Prefect, or Airflow) from Lesson 6
- An OpenTelemetry collector and a metrics backend (from Lesson 8)
- A secrets manager (Vault, AWS Secrets Manager, or cloud equivalent)
- Python 3.11+ with the Anthropic SDK installed
Your prototype likely ran the orchestrator and every agent in one process, on one machine, with a single API key and only the retries you added by hand. Production is a different system: it must survive a worker that hangs mid-generation, an endpoint that returns a 529, a deploy that ships a bad prompt, and a 3am page with no human who understands the agent graph.
The shift is from "does it work once" to "does it keep working under load, partial failure, and change." Four properties matter more than feature count:
| Property | Prototype reality | Production requirement |
|---|---|---|
| Availability | Works when you run it | SLO with an error budget |
| Recoverability | Restart and pray | Durable state, replayable runs |
| Isolation | One process, one key | Independent scaling, scoped credentials |
| Observability | print() and logs | Distributed traces and alerts |
The single most important architectural decision is to separate durable orchestration state from agent compute. The orchestrator must never hold the only copy of "where are we in this run." That belongs in a workflow engine or database so any worker can crash without losing the run. Agents are just unusually expensive, slow, non-deterministic workers — the distributed-systems rules still apply.
Different agents have different cost and latency profiles. A coordinator does cheap routing; a synthesis agent burns large context windows; a fact-checker fans out into many small calls. Share one deployment and you scale the cheap thing to feed the expensive thing. Deploy each agent role as its own service so you can size and scale them independently. A queue between orchestrator and workers gives you back-pressure, turning the model API's rate limits into queue depth instead of a cascade of failures.
┌──────────────┐
request ─────▶│ Orchestrator │ (durable state in workflow engine)
└──────┬───────┘
│ enqueue tasks
┌─────────▼──────────┐
│ task queue (SQS │ ← back-pressure lives here
│ / Redis / NATS) │
└───┬─────┬──────┬───┘
┌─────────▼─┐ ┌─▼──────┐ ┌▼────────────┐
│ search │ │ synth │ │ fact-check │ ← scaled separately
│ agent x6 │ │ agent │ │ agent x3 │
│ (cheap) │ │ x2 │ │ │
└───────────┘ └────────┘ └─────────────┘
Build one container image and select the role at runtime via an AGENT_ROLE environment variable, so you ship a single artifact and configure topology at deploy time. On Kubernetes, give each role its own Deployment and autoscale on queue depth, not CPU — agent workers are I/O-bound on the model API, so CPU autoscaling lags badly.
# search-agent.deployment.yaml (excerpt)
spec:
replicas: 6 # tuned per role, autoscaled by KEDA
template:
spec:
containers:
- name: worker
image: registry.example.com/agents:1.8.0
env:
- { name: AGENT_ROLE, value: "search" } # one image, role via env
- name: ANTHROPIC_API_KEY # per-role key, not shared
valueFrom: { secretKeyRef: { name: search-creds, key: api-key } }
You cannot promise "the agent always answers correctly" — the system is non-deterministic and the model is fallible. Define SLOs on measurable, externally meaningful properties, not on correctness you can't observe in real time. Good SLI candidates:
| SLI | Definition | Target |
|---|---|---|
| Availability | Runs completing without an unrecoverable error | 99.5% |
| Latency | End-to-end run time at p95 | < 90s |
| Task success | Runs passing automated validation (eval gates) | 95% |
| Escalation rate | Runs that fall back to a human | < 5% |
| Cost per run | p95 token spend per completed run | in budget |
Turn these into an error budget with burn-rate alerts: a fast burn (roughly 14x over 1h) pages on-call, a slow burn (6x over 6h) opens a ticket. The budget is permission to take risk — burning it fast freezes risky deploys; budget to spare means ship faster.
The reliability primitive that makes any of this safe is idempotency. Retries are guaranteed in a non-deterministic system, so every task carries an idempotency key and side-effecting tools dedupe on it — otherwise a retried "send email" agent sends the email twice.
def run_task(task):
key = task["idempotency_key"] # stable per logical task
if (cached := results.get(key)): # durable store, not in-memory
return cached # replay-safe: never re-execute
out = agent.invoke(task)
results.put(key, out, ttl=DAYS_7)
return out
A multi-agent system has more failure modes than a single service: a worker hangs, the queue backs up, an endpoint degrades, a tool returns garbage, the context grows until a call is rejected. Find these in a controlled experiment, not in the incident. Chaos testing injects failures deliberately and verifies the system degrades instead of collapses. Inject faults at the seams you own — the tool layer and the queue:
# chaos.py — wrap tool calls with controlled fault injection
import random, time
class ChaosWrapper:
def __init__(self, tool, fail_rate=0.0, latency_ms=0):
self.tool, self.fail_rate, self.latency_ms = tool, fail_rate, latency_ms
def __call__(self, *args, **kwargs):
if random.random() < self.fail_rate:
raise TimeoutError("chaos: injected tool failure")
if self.latency_ms:
time.sleep(self.latency_ms / 1000)
return self.tool(*args, **kwargs)
# Staging only: search_tool = ChaosWrapper(search_tool, fail_rate=0.2, latency_ms=800)
Run each experiment with an explicit hypothesis and a blast-radius limit — staging first, then a small traffic slice. Common experiments:
| Experiment | Inject | Expected graceful behaviour |
|---|---|---|
| Worker death | Kill a pod mid-run | Run resumes on another worker from durable state |
| Slow model | +5s latency on one role | Timeout fires, circuit breaker opens, fallback path used |
| Tool outage | 100% failure on one tool | Agent skips optional tool, marks result low-confidence |
| Queue flood | 10x enqueue rate | Back-pressure holds, autoscaler adds workers, no data loss |
Graceful degradation means defining, ahead of time, what a partial answer looks like. If the fact-check agent is down, return the synthesis with a verified: false flag rather than failing the run. Build this on the circuit breaker and fallback patterns from Lesson 7.
async def orchestrate(task):
draft = await synth_agent.run(task)
try:
checked = await with_timeout(factcheck_agent.run(draft), seconds=20)
return {"answer": checked, "verified": True}
except (TimeoutError, CircuitOpen):
DEGRADED.inc() # metric -> alert if sustained
return {"answer": draft, "verified": False, "degraded": True}
Once agents run as separate services, the messages between them cross a network. Treat agent-to-agent calls like any service-to-service traffic: authenticated, authorized, least-privilege. The principal hierarchy from the Agentic AI safety lesson still applies — but now there are many actors, and one confused agent can drive others.
Authentication. Every message carries a verifiable identity: mutual TLS between services plus short-lived signed tokens scoped to the calling agent's role. Verify the signature and role before acting on any task.
# auth.py — verify a signed agent-to-agent envelope
import jwt # short-lived per-role tokens, rotated by the secrets manager
def verify_envelope(token: str, expected_audience: str) -> dict:
claims = jwt.decode(
token, PUBLIC_KEY, algorithms=["RS256"],
audience=expected_audience, # this worker's role
)
if claims["role"] not in ALLOWED_CALLERS[expected_audience]:
raise PermissionError(f"{claims['role']} may not call {expected_audience}")
return claims
Least-privilege tools. An agent holds only the tools its job requires: the search agent gets read-only web access, only the publishing agent writes to the CMS, no agent gets broad shell access. Encode this as a per-role allowlist enforced at the tool dispatcher — never by trusting the model to stay in bounds.
TOOL_GRANTS = {
"search": {"web_search", "read_docs"},
"synth": {"read_docs"},
"factcheck": {"web_search", "read_docs"},
"publish": {"cms_write"}, # only this role can write
}
def dispatch_tool(role: str, name: str, args: dict):
if name not in TOOL_GRANTS[role]:
AUDIT.log(role=role, tool=name, decision="DENIED", args=args)
raise PermissionError(f"{role} not granted {name}")
AUDIT.log(role=role, tool=name, decision="ALLOWED", args=args)
return TOOLS[name](**args)
Audit logging. Log every tool call, handoff, and decision with run ID, agent role, and inputs — append-only, tamper-evident, correlated with the traces from Lesson 8. When something goes wrong (or a regulator asks), you must reconstruct exactly which agent did what and why.
In an agent system, your "code" includes prompts, tool definitions, and model selection — all of which change behaviour without changing a line of Python. Version everything that affects output and roll it out behind the same safeguards you use for code: canaries, gradual ramp, automatic rollback on SLO regression. Pin prompts and models as deployable config, not hardcoded strings, so you can roll them back independently of the binary.
# config — versioned and shipped as artifacts
AGENT_CONFIG = {
"search": {
"prompt_version": "search-v7", # stored in a prompt registry
"model": "claude-sonnet-latest", # name an alias, not a frozen id
"max_tokens": 4096,
"tool_grants": ["web_search", "read_docs"],
},
}
Wrap the agent step in a workflow engine activity (the Temporal @activity.defn decorator, a Prefect task, or an Airflow operator) so retries, timeouts, and heartbeats are managed durably — a non-deterministic agent call becomes a replayable step that survives worker restarts, and the activity is the unit you version. The activity loads its config at call time, so a canary can serve a new prompt_version to a slice of traffic. Roll out a prompt or model change like a canary deploy:
| Stage | Traffic | Gate to advance |
|---|---|---|
| Shadow | 0% (compare only) | Offline eval parity vs current |
| Canary | 5% | No SLO regression, cost within budget for 30 min |
| Ramp | 25% → 50% | Escalation and degraded rates stable |
| Full | 100% | Burn rate normal for 1h |
The last mile is operational. When the system pages at 3am, the responder needs a runbook, not a research project. Write one per high-probability failure, each with a symptom, a diagnosis path, and concrete mitigation steps.
RUNBOOK: degraded-response rate > 10% for 5 min
SYMPTOM alert "agent.degraded.rate high"; users get verified:false answers
DIAGNOSE 1. Check which role is degrading (dashboard: per-role error %)
2. Check model provider status page + 5xx/529 rate
3. Check circuit-breaker state for that role
MITIGATE a. If provider degraded: pin to fallback model alias
b. If one role: scale that role, drain bad pods
c. If queue backed up: confirm KEDA scaled; raise max replicas
ROLLBACK If correlated with last deploy: flip prompt_version to previous
ESCALATE If unrecovered in 15 min, page service owner; freeze deploys
Before you call the system production-ready, walk this checklist. It is the deliverable for this lesson — adapt it and keep it in the repo.
| Category | Check |
|---|---|
| Scaling | Each role scales independently on queue depth; rate limits map to back-pressure |
| Reliability | Durable run state survives any worker death; side-effecting tasks are idempotent; SLOs wired to burn-rate alerts |
| Resilience | Circuit breakers and fallbacks for every dependency; degradation emits metrics; chaos game day passed |
| Access control | mTLS plus signed role tokens; per-role tool allowlist at the dispatcher; append-only audit log; per-role secrets, rotated |
| Observability | Distributed tracing across agents (L8); cost per run budgeted (L9) |
| Rollout | Prompts and models versioned and canaried; automatic rollback on SLO regression |
| Ops | A runbook per top failure mode; on-call rotation, escalation policy, and dashboards live |
Questions & Answers
Key Takeaways
- Externalize durable state. The orchestrator is a thin scheduler over state held in a workflow engine or database. If a dying worker loses progress, you have a prototype, not a production system.
- Scale agent roles independently on queue depth. Different agents have different cost and latency profiles; give each its own deployment, its own key, and back-pressure via a queue so rate limits never cascade.
- Write SLOs you can measure, backed by error budgets. Govern availability, latency, success rate, escalation, and cost — not model correctness, which belongs to offline evals. Idempotency makes the inevitable retries safe.
- Engineer for partial failure. Chaos-test the seams, define what a degraded answer looks like before you need it, and make every fallback emit a metric so silent degradation can't hide a dying dependency.
- Authenticate and least-privilege every agent. mTLS plus signed role tokens, a per-role tool allowlist enforced at the dispatcher, fetched content treated as data rather than instructions, and an append-only audit log correlated with traces.
- Ship prompts and models like code. Version them, canary them, gate on SLOs, and keep rollback one config flip away — then run the readiness checklist as an explicit go/no-go before every launch.
Next Steps: Continue with Local LLMs