Monitoring & Observability
Learning Outcomes
- Instrument a multi-agent system with OpenTelemetry-compatible distributed tracing
- Correlate actions across agents using a single shared trace context
- Emit structured logs that capture decisions, tool calls, and reasoning
- Collect the metrics that matter — latency, token usage, and success rate per agent
- Build replay and alerting around the non-deterministic failure modes of agent fleets
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why observability is harder for agents than for services |
| Explain | 8 min | The three pillars and the GenAI semantic conventions |
| Demo | 9 min | Distributed tracing with context propagation |
| Demo | 8 min | Structured logging of decisions and tool calls |
| Demo | 7 min | Metrics: latency, tokens, success rate per agent |
| Demo | 7 min | Replay, SLOs, alerting, and the collector |
| Wrap-up | 3 min | Key takeaways and next steps |
Before You Begin
Pre-work:
- Complete Lesson 7: Error Handling & Recovery — observability tells your retry and circuit-breaker logic when to fire.
- Have a working multi-agent system from Lesson 5: The Anthropic Agent SDK or Lesson 6: Workflow Engines.
- Review agent evaluation from the Agentic AI course — eval and runtime metrics share a vocabulary.
Shopping List:
- Python 3.10+ with
opentelemetry-sdk,opentelemetry-exporter-otlp, and theanthropicSDK - A running OpenTelemetry Collector (or Jaeger/Tempo/Honeycomb endpoint) on
localhost:4317 dockerfor the local collector and visualization stack
The three pillars — traces, logs, metrics — assume a deterministic request mapping to one call tree, errors that throw, and fixed per-request cost. Multi-agent systems break all of that: reasoning is non-deterministic, a goal fans out to N agents and sub-agents, cost varies 10x by how much an agent decided to think, and the worst "errors" are confident wrong answers returning HTTP 200. A supervisor that delegates to the wrong specialist produces a perfectly successful-looking trace and a useless result. Observability for agents must capture not just what happened but what the agent decided and why.
Don't invent attribute names. OpenTelemetry publishes GenAI semantic conventions that standardise span and attribute names for LLMs and agents, keeping traces portable across Datadog, Honeycomb, Langfuse, and Grafana. They are still marked Development (experimental), so pin a version and gate transitions with OTEL_SEMCONV_STABILITY_OPT_IN. They define four agent span names (create_agent, invoke_agent INTERNAL/CLIENT, invoke_workflow) plus a core attribute set:
SPAN_ATTRS = {
"gen_ai.operation.name": "invoke_agent", # required
"gen_ai.provider.name": "anthropic", # required
"gen_ai.agent.name": "fact_checker", # + gen_ai.agent.id
"gen_ai.request.model": "claude-sonnet", # model invoked
"gen_ai.usage.input_tokens": 1842, # + output_tokens (recommended)
}
Distributed tracing gives you one trace_id that stitches the supervisor, every worker, and every tool call into a single tree, via W3C trace context propagation: the parent injects a traceparent, the child extracts it and parents its spans under it.
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
provider = TracerProvider()
provider.add_span_processor(
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://localhost:4317")))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("orchestrator")
def supervisor(goal: str):
with tracer.start_as_current_span("invoke_workflow") as span:
span.set_attribute("gen_ai.operation.name", "invoke_workflow")
plan = decompose(goal) # child spans inherit this context
return synthesize([run_agent(a, t) for a, t in plan])
def run_agent(name: str, task: str):
with tracer.start_as_current_span("invoke_agent") as span:
span.set_attribute("gen_ai.agent.name", name)
return agent_call(name, task)
When agents run in separate processes behind a queue, the producer serialises context onto the message; the worker restores it before opening its span:
from opentelemetry.propagate import inject, extract
carrier = {}
inject(carrier) # producer
queue.publish({"task": task, "trace_context": carrier})
ctx = extract(queue.consume()["trace_context"]) # worker
with tracer.start_as_current_span("invoke_agent", context=ctx):
handle(task)
The Anthropic Agent SDK does this for you: in Agent SDK and non-interactive claude -p sessions, Claude Code reads TRACEPARENT and TRACESTATE from its environment and makes its claude_code.interaction span a child of your caller span, so an SDK sub-agent slots straight into your trace tree.
Traces tell you the shape; logs tell you the substance. The high-value events are decisions, tool calls, and handoffs. Emit them as structured JSON keyed by the same trace_id so logs and traces join in your backend.
import json, logging
from opentelemetry import trace
logger = logging.getLogger("agent")
def log_decision(event: str, **fields):
ctx = trace.get_current_span().get_span_context()
logger.info(json.dumps({
"event": event,
"trace_id": format(ctx.trace_id, "032x"),
"span_id": format(ctx.span_id, "016x"),
**fields,
}))
# In the supervisor, log the routing decision and its rationale
log_decision("route", agent="data_analyst", confidence=0.82,
reason="task mentions CSV aggregation; analyst owns pandas tools")
The Anthropic Agent SDK exposes this same data through PreToolUse and PostToolUse hooks, which run in-process and capture tool inputs, outputs, and timing. Content capture is off by default and gated: OTEL_LOG_TOOL_DETAILS records tool parameters, OTEL_LOG_TOOL_CONTENT records inputs and outputs (truncated at 60 KB), and raw API bodies need an explicit opt-in.
A disciplined log taxonomy pays for itself:
| Event | Fields | Answers |
|---|---|---|
route |
agent, reason, confidence | Why this worker? |
tool_call |
name, args_schema, duration_ms, ok | Which tool, how long, did it succeed? |
handoff |
from_agent, to_agent, context_bytes | Where did control move, how much state? |
retry |
attempt, error_class, backoff_ms | Recovering or looping? |
escalate |
reason, to=human | When did we ask a person? |
Traces and logs are per-request; metrics are the aggregates that drive dashboards and alerts. The three that matter for an agent fleet are latency, token usage, and success rate, all dimensioned per agent so you can see which specialist is slow, expensive, or unreliable.
from opentelemetry import metrics
meter = metrics.get_meter("orchestrator")
latency = meter.create_histogram("agent.invocation.duration", unit="s")
tokens = meter.create_counter("agent.tokens", unit="token")
outcomes = meter.create_counter("agent.outcomes")
def record(agent, model, seconds, usage, ok):
attrs = {"gen_ai.agent.name": agent, "gen_ai.request.model": model}
latency.record(seconds, attrs)
tokens.add(usage.input_tokens, {**attrs, "type": "input"})
tokens.add(usage.output_tokens, {**attrs, "type": "output"})
outcomes.add(1, {**attrs, "outcome": "ok" if ok else "fail"})
Claude Code and the Agent SDK export equivalents natively over OTLP. Set CLAUDE_CODE_ENABLE_TELEMETRY=1, OTEL_METRICS_EXPORTER=otlp, and OTEL_EXPORTER_OTLP_ENDPOINT to your collector. Standard metric names include claude_code.token.usage (broken down by type, model, and agent.name), claude_code.cost.usage in USD, claude_code.session.count, and claude_code.active_time.total.
OTEL_METRICS_INCLUDE_SESSION_ID for this reason — keep session_id and user_id on traces and logs, not on metrics.api.error.rate with a falling task.success.rate is the signature of prompt or tool drift.Because agents are non-deterministic, you cannot reliably reproduce a failure by re-running it. Instead, capture enough state to replay the exact sequence offline — per span, the prompt, tool inputs/outputs, and model parameters, keyed by trace_id — so a failed production run becomes a deterministic fixture.
import json, time
def capture_step(trace_id, step):
"""Append one replayable step to a per-trace event log."""
with open(f"replays/{trace_id}.jsonl", "a") as f:
f.write(json.dumps({
"ts": time.time(),
"agent": step["agent"], "model": step["model"],
"prompt": step["prompt"], # redact before persisting
"tool_calls": step["tool_calls"], # name, args, result
"output": step["output"],
}) + "\n")
To replay, read the JSONL back and feed the recorded tool_calls results in instead of hitting live tools, so the run is deterministic. Workflow engines give you this for free — Temporal's event-sourced history replays a workflow's exact execution against new code, invaluable when the "code" is a prompt you just changed and you want to know whether your fix would have rescued yesterday's failed runs.
Page on symptoms users feel, not every model hiccup. Define SLOs over task success rate, end-to-end latency, and cost per task, and alert on burn rate.
| SLI | Example SLO | Alert signal |
|---|---|---|
| Task success rate | 99% of goals complete without escalation | Burn rate over 1h + 6h windows |
| End-to-end latency | p95 under 90s for standard goals | p95 breaching for 10 min |
| Cost per task | p95 under a fixed USD ceiling | 2x jump in claude_code.cost.usage |
| Loop / runaway | < 0.1% of runs exceed N hops | Trace depth over threshold |
Two alert conditions are unique to agents. Runaway loops — alert when a trace exceeds a hop or span-depth budget, because an agent stuck re-calling a tool drains your token budget while every call returns 200. Silent quality drift — alert when task success rate falls while API error rate stays flat.
MAX_HOPS = 25
def guard_hops(span):
hops = span.attributes.get("orchestration.hop_count", 0)
if hops > MAX_HOPS:
emit_alert("runaway_loop", trace_id=span.context.trace_id, hops=hops)
raise RunawayLoopError(f"{hops} hops exceeds budget {MAX_HOPS}")
Finally, send everything to an OpenTelemetry Collector rather than pointing your app at a backend directly. The collector decouples you from vendor choice, batches and retries exports, and centralises sampling and PII redaction:
# otel-collector.yaml
receivers:
otlp:
protocols: { grpc: { endpoint: 0.0.0.0:4317 } }
processors:
batch: {}
attributes/redact:
actions: [ { key: gen_ai.prompt, action: delete } ] # strip prompts
exporters:
otlp/traces: { endpoint: tempo:4317, tls: { insecure: true } }
prometheus: { endpoint: 0.0.0.0:8889 }
service:
pipelines:
traces: { receivers: [otlp], processors: [batch, attributes/redact], exporters: [otlp/traces] }
metrics: { receivers: [otlp], processors: [batch], exporters: [prometheus] }
Point Tempo/Jaeger at the traces pipeline for the agent-tree waterfall and Grafana at the Prometheus endpoint for per-agent dashboards. The waterfall is your primary debugging surface — one glance shows fan-out width, the slowest worker, and whether a sub-agent is recursing.
Questions & Answers
OTEL_LOG_TOOL_CONTENT for this reason. Redact at the collector with an attributes processor so it cannot be bypassed by one misconfigured service, and enable full-content capture only for scoped, time-boxed debugging.inject() to serialize the active context onto the message; the worker calls extract() to restore it before opening its span. The Agent SDK does this automatically by reading TRACEPARENT and TRACESTATE from the environment in non-interactive sessions.Key Takeaways
- Agents need decision-level observability — capture the prompt, tool inputs, reasoning, and chosen branch, not just timing, because bad decisions look like successful requests.
- Use the GenAI semantic conventions — standard span names (
invoke_agent,invoke_workflow) andgen_ai.*attributes keep your telemetry portable, but pin the version since they are still in Development. - Propagate one trace across the whole fleet — W3C trace context via inject/extract (and the Agent SDK's
TRACEPARENThandling) turns a scattered set of calls into one debuggable tree. - Dimension metrics per agent, control cardinality — latency, tokens, and success rate per agent drive your dashboards; keep high-cardinality labels off metrics and on traces.
- Plan for replay and drift, not just exceptions — record replayable state for sampled and failed runs, and alert on runaway loops and falling task success rate, which exceptions will never reveal.
Next Steps: Lesson 9: Cost Management & Optimization