Cost Management & Optimization
Learning Outcomes
- Budget token allowances across agents and enforce them at the orchestration layer
- Apply prompt caching to eliminate redundant context costs in repeated agent calls
- Route tasks to the cheapest model that can do the job using a complexity classifier
- Cache deterministic intermediate results to avoid paying for identical work twice
- Instrument a cost-aware orchestrator that reports spend in real time and trips a kill switch
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why multi-agent systems leak |
| Explain | 6 min | Cost profiles of orchestration |
| Demo | 7 min | Token budgeting per agent |
| Demo | 8 min | Prompt caching across agent calls |
| Demo | 7 min | Intelligent model routing |
| Demo | 6 min | Result caching with idempotency keys |
| Wrap-up | 3 min | Cost-aware orchestrator, takeaways |
Before You Begin
Pre-work:
- Complete Lesson 5: The Anthropic Agent SDK and Lesson 8: Monitoring & Observability — you need a working multi-agent system to instrument
- Refresh tool use and the agent loop from Agentic AI if needed
- Have billing or
/cost-style telemetry wired up
Shopping List:
- Python 3.10+ with
pip install anthropic claude-agent-sdk - An
ANTHROPIC_API_KEYin your environment - Redis (or any key-value store) for the cache
Single-agent cost is roughly linear; orchestrated systems are not. Shared context — system prompt, task spec, retrieved docs — is re-sent to every agent. A supervisor fanning out to five workers, each with 8K tokens of briefing, pays for it six times before any work happens. Fan-out multiplies the input bill; long histories multiply again. Four levers, by impact:
| Lever | Mechanism | Where |
|---|---|---|
| Prompt caching | Stop re-billing static context | Step 3 |
| Intelligent routing | Easy tasks to Haiku, not Opus | Step 4 |
| Result caching | Never repeat a deterministic sub-task | Step 5 |
| Batching | Async work at 50% off, stacks with caching | Step 6 |
Output costs ~5x input across the tiers, so trimming a verbose agent's output often beats trimming the prompt.
Treat tokens like any constrained resource: budget per agent, decrement on each call, fail loudly on overrun. This turns "the bill was huge" into a bounded SLO.
class BudgetExceeded(Exception): pass
class TokenBudget: # per-agent ceiling
def __init__(self, limit): self.limit, self.spent = limit, 0
def charge(self, in_tok, out_tok):
weighted = in_tok + out_tok * 5 # output is ~5x input
if self.spent + weighted > self.limit:
raise BudgetExceeded(f"{self.spent + weighted} > {self.limit}")
self.spent += weighted
budgets = {"coordinator": TokenBudget(50_000), "search": TokenBudget(200_000),
"synthesis": TokenBudget(120_000), "fact_checker": TokenBudget(80_000)}
Charge real usage from each response (resp.usage carries input_tokens, output_tokens, and cache fields); Step 6 wires charge(...) into the run loop. On overrun, escalate (Lesson 7), downgrade, or truncate. Budgets work best per-run and per-agent so a runaway worker never drains it.
max_tokens deliberately per agent, not a copy-pasted 4096. An agent returning a one-word label should cap output near 16 tokens. Output dominates cost, so this is the cheapest guardrail.Prompt caching is the highest-leverage optimization here because so much context is shared and stable. Mark a stable prefix with cache_control; the first call writes the cache, later calls within the TTL read it cheaply. Multipliers:
| Token type | Multiplier vs base input |
|---|---|
| 5-min cache write | 1.25x |
| 1-hour cache write | 2x |
| Cache read | 0.1x |
A read costs one-tenth of fresh input, so caching pays off the moment a prefix is reused even twice. Build system as a list of text blocks and attach an ephemeral cache_control to the last stable one — Anthropic caches everything up to and including it, leaving the volatile messages uncached. For a longer window, set that block's ttl to 1 hour. On a hit, usage reports most input under cache_read_input_tokens.
Not every sub-task needs your most capable model. The Claude lineup spans tiers — Haiku (low-cost), Sonnet (middle), Opus (premium). Routing classification to Haiku not Opus is a step-change in cost. The classifier that picks the tier must itself be cheap — smallest model, capped output, tiny prompt:
TIERS = {"trivial": "claude-haiku-4-5", "standard": "claude-sonnet-4-6",
"hard": "claude-opus-4-8"}
def route(task: str) -> str: # one cheap call picks a model
r = client.messages.create(model="claude-haiku-4-5", max_tokens=8,
system="Reply with one word: trivial, standard, or hard.",
messages=[{"role": "user", "content": task[:2000]}])
return TIERS.get(r.content[0].text.strip().lower(), "claude-sonnet-4-6")
In the Agent SDK you often need no separate router — each subagent sets its own model in its AgentDefinition (a "triager" on Haiku, an "analyst" on Opus). Subagents run in isolated context windows and report only conclusions back, so the top-level agent never carries the triager's history.
Prompt caching saves on input; result caching saves on the entire call. If a deterministic sub-task ran before with identical inputs, return the stored answer — ten workers looking up the same entity become one call. Key on a stable hash of model, prompt version, and inputs (cache_key() in Step 6); a PROMPT_VERSION tag invalidates it cleanly.
Cache only deterministic, side-effect-free steps — extraction, classification, validation. Never cache time-sensitive lookups, side-effecting steps, per-user output, or anything where stale is unsafe.
Compose the pieces behind one entry: route, cache, charge, bound.
import hashlib, json, redis
cache = redis.Redis(decode_responses=True)
PROMPT_VERSION = "v3"
spent = {"usd": 0.0} # shared cost meter; emit to your metrics pipeline
def cache_key(model, task, inputs): # step 5
p = json.dumps({"m": model, "v": PROMPT_VERSION, "t": task, "i": inputs},
sort_keys=True)
return "agentcache:" + hashlib.sha256(p.encode()).hexdigest()
def run(agent, task, inputs, limit_usd):
if spent["usd"] >= limit_usd:
raise BudgetExceeded("global run budget exhausted") # kill switch
key = cache_key(model := route(task), task, inputs) # steps 4 + 5
if (hit := cache.get(key)) is not None:
return json.loads(hit) # cached: zero tokens billed
u = (r := client.messages.create(model=model, max_tokens=2048,
messages=[{"role": "user", "content": task}])).usage
budgets[agent].charge(u.input_tokens, u.output_tokens) # step 2
spent["usd"] += usd_cost(model, u) # cache reads billed at 0.1x
cache.set(key, json.dumps(out := r.content[0].text), ex=3600)
return out
usd_cost applies per-model rates from config (cache reads at 0.1x) — keep them out of code because they move. The limit_usd check is a circuit breaker; pair it with Lesson 7's recovery patterns and emit spent to your Lesson 8 telemetry.
Questions & Answers
Key Takeaways
- Orchestration multiplies cost — Shared context is re-billed to every agent; fix topology first.
- Prompt caching is the biggest lever — Reads cost one-tenth of input; cache stable prefixes, keep volatile bytes after.
- Budget per agent, per run — Charge real
usageagainst a weighted ceiling (output ~5x) so runaways fail fast. - Route to the cheapest capable model — A classifier or per-subagent model turns Opus calls into Haiku; watch escalation rate.
- Result-cache deterministic steps only — Use a versioned idempotency key; a stale entry is a bug.
- Make cost observable and bounded — Emit per-agent spend to the Lesson 8 pipeline; add a USD kill switch.
Next Steps: Lesson 10: Production Multi-Agent Systems