Cost Management & Optimization

40 min advanced Lesson 9

Learning Outcomes

  • Budget token allowances across agents and enforce them at the orchestration layer
  • Apply prompt caching to eliminate redundant context costs in repeated agent calls
  • Route tasks to the cheapest model that can do the job using a complexity classifier
  • Cache deterministic intermediate results to avoid paying for identical work twice
  • Instrument a cost-aware orchestrator that reports spend in real time and trips a kill switch

Lesson Plan

Segment Duration Topic
Intro 3 min Why multi-agent systems leak
Explain 6 min Cost profiles of orchestration
Demo 7 min Token budgeting per agent
Demo 8 min Prompt caching across agent calls
Demo 7 min Intelligent model routing
Demo 6 min Result caching with idempotency keys
Wrap-up 3 min Cost-aware orchestrator, takeaways

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ with pip install anthropic claude-agent-sdk
  • An ANTHROPIC_API_KEY in your environment
  • Redis (or any key-value store) for the cache

1 Where Multi-Agent Cost Actually Comes From

Single-agent cost is roughly linear; orchestrated systems are not. Shared context — system prompt, task spec, retrieved docs — is re-sent to every agent. A supervisor fanning out to five workers, each with 8K tokens of briefing, pays for it six times before any work happens. Fan-out multiplies the input bill; long histories multiply again. Four levers, by impact:

Lever Mechanism Where
Prompt caching Stop re-billing static context Step 3
Intelligent routing Easy tasks to Haiku, not Opus Step 4
Result caching Never repeat a deterministic sub-task Step 5
Batching Async work at 50% off, stacks with caching Step 6

Output costs ~5x input across the tiers, so trimming a verbose agent's output often beats trimming the prompt.

WARNING
Watch Out
Pattern choice is itself a cost decision. A hierarchy with three supervisor layers re-broadcasts context at every level. Before optimizing token-by-token, ask whether a flatter pipeline (Lesson 2) carries less of it.

2 Token Budgeting Across Agents

Treat tokens like any constrained resource: budget per agent, decrement on each call, fail loudly on overrun. This turns "the bill was huge" into a bounded SLO.

class BudgetExceeded(Exception): pass

class TokenBudget:                          # per-agent ceiling
    def __init__(self, limit): self.limit, self.spent = limit, 0
    def charge(self, in_tok, out_tok):
        weighted = in_tok + out_tok * 5     # output is ~5x input
        if self.spent + weighted > self.limit:
            raise BudgetExceeded(f"{self.spent + weighted} > {self.limit}")
        self.spent += weighted

budgets = {"coordinator": TokenBudget(50_000), "search": TokenBudget(200_000),
           "synthesis": TokenBudget(120_000), "fact_checker": TokenBudget(80_000)}

Charge real usage from each response (resp.usage carries input_tokens, output_tokens, and cache fields); Step 6 wires charge(...) into the run loop. On overrun, escalate (Lesson 7), downgrade, or truncate. Budgets work best per-run and per-agent so a runaway worker never drains it.

TIP
Tip
Set max_tokens deliberately per agent, not a copy-pasted 4096. An agent returning a one-word label should cap output near 16 tokens. Output dominates cost, so this is the cheapest guardrail.

3 Prompt Caching: Stop Paying for Static Context

Prompt caching is the highest-leverage optimization here because so much context is shared and stable. Mark a stable prefix with cache_control; the first call writes the cache, later calls within the TTL read it cheaply. Multipliers:

Token type Multiplier vs base input
5-min cache write 1.25x
1-hour cache write 2x
Cache read 0.1x

A read costs one-tenth of fresh input, so caching pays off the moment a prefix is reused even twice. Build system as a list of text blocks and attach an ephemeral cache_control to the last stable one — Anthropic caches everything up to and including it, leaving the volatile messages uncached. For a longer window, set that block's ttl to 1 hour. On a hit, usage reports most input under cache_read_input_tokens.

WARNING
Watch Out
Caching is prefix-based and order-sensitive (tools, system, messages). Change any earlier byte and the cache misses, re-charging the write premium — keep volatile content (timestamps, IDs, the task) after cached blocks. There is also a minimum cacheable length: short prompts run uncached with no error, so reserve it for large, stable material.

4 Intelligent Routing: Match the Model to the Task

Not every sub-task needs your most capable model. The Claude lineup spans tiers — Haiku (low-cost), Sonnet (middle), Opus (premium). Routing classification to Haiku not Opus is a step-change in cost. The classifier that picks the tier must itself be cheap — smallest model, capped output, tiny prompt:

TIERS = {"trivial": "claude-haiku-4-5", "standard": "claude-sonnet-4-6",
         "hard": "claude-opus-4-8"}

def route(task: str) -> str:                       # one cheap call picks a model
    r = client.messages.create(model="claude-haiku-4-5", max_tokens=8,
        system="Reply with one word: trivial, standard, or hard.",
        messages=[{"role": "user", "content": task[:2000]}])
    return TIERS.get(r.content[0].text.strip().lower(), "claude-sonnet-4-6")

In the Agent SDK you often need no separate router — each subagent sets its own model in its AgentDefinition (a "triager" on Haiku, an "analyst" on Opus). Subagents run in isolated context windows and report only conclusions back, so the top-level agent never carries the triager's history.

TIP
Tip
Measure routing accuracy, not just savings. A router that mislabels a hard task as trivial yields a wrong answer you then redo on Opus — costlier than routing right the first time. Track the escalation rate and tune until it is stable.

5 Result Caching with Idempotency Keys

Prompt caching saves on input; result caching saves on the entire call. If a deterministic sub-task ran before with identical inputs, return the stored answer — ten workers looking up the same entity become one call. Key on a stable hash of model, prompt version, and inputs (cache_key() in Step 6); a PROMPT_VERSION tag invalidates it cleanly.

Cache only deterministic, side-effect-free steps — extraction, classification, validation. Never cache time-sensitive lookups, side-effecting steps, per-user output, or anything where stale is unsafe.

WARNING
Watch Out
A result cache is a correctness surface, not just a cost one — stale or wrongly-keyed entries silently corrupt downstream agents. Key on model and prompt-version, set a defensible TTL, and track hit rate so a bad key shows up.

6 A Cost-Aware Orchestrator

Compose the pieces behind one entry: route, cache, charge, bound.

import hashlib, json, redis
cache = redis.Redis(decode_responses=True)
PROMPT_VERSION = "v3"
spent = {"usd": 0.0}   # shared cost meter; emit to your metrics pipeline

def cache_key(model, task, inputs):                # step 5
    p = json.dumps({"m": model, "v": PROMPT_VERSION, "t": task, "i": inputs},
                   sort_keys=True)
    return "agentcache:" + hashlib.sha256(p.encode()).hexdigest()

def run(agent, task, inputs, limit_usd):
    if spent["usd"] >= limit_usd:
        raise BudgetExceeded("global run budget exhausted")   # kill switch
    key = cache_key(model := route(task), task, inputs)       # steps 4 + 5
    if (hit := cache.get(key)) is not None:
        return json.loads(hit)                     # cached: zero tokens billed
    u = (r := client.messages.create(model=model, max_tokens=2048,
         messages=[{"role": "user", "content": task}])).usage
    budgets[agent].charge(u.input_tokens, u.output_tokens)    # step 2
    spent["usd"] += usd_cost(model, u)             # cache reads billed at 0.1x
    cache.set(key, json.dumps(out := r.content[0].text), ex=3600)
    return out

usd_cost applies per-model rates from config (cache reads at 0.1x) — keep them out of code because they move. The limit_usd check is a circuit breaker; pair it with Lesson 7's recovery patterns and emit spent to your Lesson 8 telemetry.

TIP
Tip
Route the largest non-urgent backlog through the Message Batches API for an automatic 50% discount on input and output — it stacks with prompt caching. If you tolerate a 24-hour SLA, it is free money.

Questions & Answers

Q: The cache TTL is 5 minutes but my fan-out takes longer — does it expire mid-run and cost the write twice?
It can. Each read refreshes the 5-minute lifetime, but a gap longer than the TTL evicts the prefix and the next call re-pays the write premium. For fan-outs spanning minutes, request the 1-hour TTL — its higher multiplier (2x vs 1.25x) is cheap insurance.
Q: Won't a routing classifier add another point of failure and latency?
It adds one cheap Haiku call (a few output tokens, often sub-200ms). The failure mode that matters is misclassification — default to a safe tier on any unparseable response and track the escalation rate; if too many "trivial" results bounce back for redo, savings evaporate. The simpler win is often per-subagent models in the Agent SDK, needing no separate router.
Q: Is result caching safe when agents have side effects?
Only cache deterministic, side-effect-free steps. Never cache a step that sends mail, writes a row, or reads a live value. Key on model, prompt version, and inputs with a defensible TTL. A wrongly-keyed cache propagates a correctness bug downstream — treat it like any shared-state component (Lesson 3).
Q: Prices and model names change constantly. How do I keep cost code from rotting?
Keep rates and model IDs in configuration, never in logic. Express optimizations in durable terms — "route to the cheapest tier that passes the classifier," "cap output since it is ~5x input" — that survive model generations. Then alert on cost-per-run drift so a silent price change is a graph anomaly, not a surprise invoice.

Key Takeaways

  1. Orchestration multiplies cost — Shared context is re-billed to every agent; fix topology first.
  2. Prompt caching is the biggest lever — Reads cost one-tenth of input; cache stable prefixes, keep volatile bytes after.
  3. Budget per agent, per run — Charge real usage against a weighted ceiling (output ~5x) so runaways fail fast.
  4. Route to the cheapest capable model — A classifier or per-subagent model turns Opus calls into Haiku; watch escalation rate.
  5. Result-cache deterministic steps only — Use a versioned idempotency key; a stale entry is a bug.
  6. Make cost observable and bounded — Emit per-agent spend to the Lesson 8 pipeline; add a USD kill switch.

Next Steps: Lesson 10: Production Multi-Agent Systems