Monitoring & Observability

45 min advanced Lesson 8

Learning Outcomes

  • Instrument a multi-agent system with OpenTelemetry-compatible distributed tracing
  • Correlate actions across agents using a single shared trace context
  • Emit structured logs that capture decisions, tool calls, and reasoning
  • Collect the metrics that matter — latency, token usage, and success rate per agent
  • Build replay and alerting around the non-deterministic failure modes of agent fleets

Lesson Plan

Segment Duration Topic
Intro 3 min Why observability is harder for agents than for services
Explain 8 min The three pillars and the GenAI semantic conventions
Demo 9 min Distributed tracing with context propagation
Demo 8 min Structured logging of decisions and tool calls
Demo 7 min Metrics: latency, tokens, success rate per agent
Demo 7 min Replay, SLOs, alerting, and the collector
Wrap-up 3 min Key takeaways and next steps

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ with opentelemetry-sdk, opentelemetry-exporter-otlp, and the anthropic SDK
  • A running OpenTelemetry Collector (or Jaeger/Tempo/Honeycomb endpoint) on localhost:4317
  • docker for the local collector and visualization stack

1 Why Agents Are Different — and the GenAI Conventions

The three pillars — traces, logs, metrics — assume a deterministic request mapping to one call tree, errors that throw, and fixed per-request cost. Multi-agent systems break all of that: reasoning is non-deterministic, a goal fans out to N agents and sub-agents, cost varies 10x by how much an agent decided to think, and the worst "errors" are confident wrong answers returning HTTP 200. A supervisor that delegates to the wrong specialist produces a perfectly successful-looking trace and a useless result. Observability for agents must capture not just what happened but what the agent decided and why.

Don't invent attribute names. OpenTelemetry publishes GenAI semantic conventions that standardise span and attribute names for LLMs and agents, keeping traces portable across Datadog, Honeycomb, Langfuse, and Grafana. They are still marked Development (experimental), so pin a version and gate transitions with OTEL_SEMCONV_STABILITY_OPT_IN. They define four agent span names (create_agent, invoke_agent INTERNAL/CLIENT, invoke_workflow) plus a core attribute set:

SPAN_ATTRS = {
    "gen_ai.operation.name": "invoke_agent",   # required
    "gen_ai.provider.name": "anthropic",       # required
    "gen_ai.agent.name": "fact_checker",       # + gen_ai.agent.id
    "gen_ai.request.model": "claude-sonnet",   # model invoked
    "gen_ai.usage.input_tokens": 1842,         # + output_tokens (recommended)
}
NOTE
Key Insight
Treat every agent decision as a span and every tool call as a child span — the shape of the trace tree is your single most valuable debugging artifact. A 40-deep nested trace usually means an agent is looping; a flat trace with one giant span usually means a context window blew up. Conventions are still in Development, so wrap attribute names behind a constants module to make a version bump a one-file change.

2 Distributed Tracing With Context Propagation

Distributed tracing gives you one trace_id that stitches the supervisor, every worker, and every tool call into a single tree, via W3C trace context propagation: the parent injects a traceparent, the child extracts it and parents its spans under it.

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter

provider = TracerProvider()
provider.add_span_processor(
    BatchSpanProcessor(OTLPSpanExporter(endpoint="http://localhost:4317")))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("orchestrator")

def supervisor(goal: str):
    with tracer.start_as_current_span("invoke_workflow") as span:
        span.set_attribute("gen_ai.operation.name", "invoke_workflow")
        plan = decompose(goal)            # child spans inherit this context
        return synthesize([run_agent(a, t) for a, t in plan])

def run_agent(name: str, task: str):
    with tracer.start_as_current_span("invoke_agent") as span:
        span.set_attribute("gen_ai.agent.name", name)
        return agent_call(name, task)

When agents run in separate processes behind a queue, the producer serialises context onto the message; the worker restores it before opening its span:

from opentelemetry.propagate import inject, extract

carrier = {}
inject(carrier)                               # producer
queue.publish({"task": task, "trace_context": carrier})

ctx = extract(queue.consume()["trace_context"])  # worker
with tracer.start_as_current_span("invoke_agent", context=ctx):
    handle(task)

The Anthropic Agent SDK does this for you: in Agent SDK and non-interactive claude -p sessions, Claude Code reads TRACEPARENT and TRACESTATE from its environment and makes its claude_code.interaction span a child of your caller span, so an SDK sub-agent slots straight into your trace tree.

TIP
Tip
Propagate context across your message queue from day one. Retrofitting trace propagation onto a shipped event-driven system means touching every producer and consumer at once — exactly the migration tracing was supposed to save you from.

3 Structured Logging of Decisions and Tool Calls

Traces tell you the shape; logs tell you the substance. The high-value events are decisions, tool calls, and handoffs. Emit them as structured JSON keyed by the same trace_id so logs and traces join in your backend.

import json, logging
from opentelemetry import trace

logger = logging.getLogger("agent")

def log_decision(event: str, **fields):
    ctx = trace.get_current_span().get_span_context()
    logger.info(json.dumps({
        "event": event,
        "trace_id": format(ctx.trace_id, "032x"),
        "span_id": format(ctx.span_id, "016x"),
        **fields,
    }))

# In the supervisor, log the routing decision and its rationale
log_decision("route", agent="data_analyst", confidence=0.82,
             reason="task mentions CSV aggregation; analyst owns pandas tools")

The Anthropic Agent SDK exposes this same data through PreToolUse and PostToolUse hooks, which run in-process and capture tool inputs, outputs, and timing. Content capture is off by default and gated: OTEL_LOG_TOOL_DETAILS records tool parameters, OTEL_LOG_TOOL_CONTENT records inputs and outputs (truncated at 60 KB), and raw API bodies need an explicit opt-in.

A disciplined log taxonomy pays for itself:

Event Fields Answers
route agent, reason, confidence Why this worker?
tool_call name, args_schema, duration_ms, ok Which tool, how long, did it succeed?
handoff from_agent, to_agent, context_bytes Where did control move, how much state?
retry attempt, error_class, backoff_ms Recovering or looping?
escalate reason, to=human When did we ask a person?
WARNING
Watch Out
Logging full prompts and tool outputs is a PII and secrets minefield — tool inputs routinely contain customer records, API keys passed as arguments, and internal URLs. Default to logging schema and lengths, redact sensitive fields, and turn on full-content capture only for a scoped debugging window, never as a standing default in production.

4 Metrics: Latency, Tokens, and Success Rate Per Agent

Traces and logs are per-request; metrics are the aggregates that drive dashboards and alerts. The three that matter for an agent fleet are latency, token usage, and success rate, all dimensioned per agent so you can see which specialist is slow, expensive, or unreliable.

from opentelemetry import metrics
meter = metrics.get_meter("orchestrator")
latency  = meter.create_histogram("agent.invocation.duration", unit="s")
tokens   = meter.create_counter("agent.tokens", unit="token")
outcomes = meter.create_counter("agent.outcomes")

def record(agent, model, seconds, usage, ok):
    attrs = {"gen_ai.agent.name": agent, "gen_ai.request.model": model}
    latency.record(seconds, attrs)
    tokens.add(usage.input_tokens, {**attrs, "type": "input"})
    tokens.add(usage.output_tokens, {**attrs, "type": "output"})
    outcomes.add(1, {**attrs, "outcome": "ok" if ok else "fail"})

Claude Code and the Agent SDK export equivalents natively over OTLP. Set CLAUDE_CODE_ENABLE_TELEMETRY=1, OTEL_METRICS_EXPORTER=otlp, and OTEL_EXPORTER_OTLP_ENDPOINT to your collector. Standard metric names include claude_code.token.usage (broken down by type, model, and agent.name), claude_code.cost.usage in USD, claude_code.session.count, and claude_code.active_time.total.

WARNING
Watch Out
Per-agent, per-model, per-session labels multiply into a cardinality explosion that can bankrupt a metrics backend faster than your token bill. Claude Code ships controls such as OTEL_METRICS_INCLUDE_SESSION_ID for this reason — keep session_id and user_id on traces and logs, not on metrics.
NOTE
Key Insight
Define success rate at the orchestration layer, not the API layer. A 200 from the model is not success — success is the supervisor accepting the worker's output. A healthy api.error.rate with a falling task.success.rate is the signature of prompt or tool drift.

5 Replay and Debugging Non-Deterministic Runs

Because agents are non-deterministic, you cannot reliably reproduce a failure by re-running it. Instead, capture enough state to replay the exact sequence offline — per span, the prompt, tool inputs/outputs, and model parameters, keyed by trace_id — so a failed production run becomes a deterministic fixture.

import json, time

def capture_step(trace_id, step):
    """Append one replayable step to a per-trace event log."""
    with open(f"replays/{trace_id}.jsonl", "a") as f:
        f.write(json.dumps({
            "ts": time.time(),
            "agent": step["agent"], "model": step["model"],
            "prompt": step["prompt"],          # redact before persisting
            "tool_calls": step["tool_calls"],  # name, args, result
            "output": step["output"],
        }) + "\n")

To replay, read the JSONL back and feed the recorded tool_calls results in instead of hitting live tools, so the run is deterministic. Workflow engines give you this for free — Temporal's event-sourced history replays a workflow's exact execution against new code, invaluable when the "code" is a prompt you just changed and you want to know whether your fix would have rescued yesterday's failed runs.

TIP
Tip
Sample full-content capture: persist complete replay payloads for a small percentage of normal runs plus 100% of runs that error or escalate. You get cheap forensic coverage of failures without paying to store the prompts of every successful request.

6 SLOs, Alerting, and Wiring It Together

Page on symptoms users feel, not every model hiccup. Define SLOs over task success rate, end-to-end latency, and cost per task, and alert on burn rate.

SLI Example SLO Alert signal
Task success rate 99% of goals complete without escalation Burn rate over 1h + 6h windows
End-to-end latency p95 under 90s for standard goals p95 breaching for 10 min
Cost per task p95 under a fixed USD ceiling 2x jump in claude_code.cost.usage
Loop / runaway < 0.1% of runs exceed N hops Trace depth over threshold

Two alert conditions are unique to agents. Runaway loops — alert when a trace exceeds a hop or span-depth budget, because an agent stuck re-calling a tool drains your token budget while every call returns 200. Silent quality drift — alert when task success rate falls while API error rate stays flat.

MAX_HOPS = 25

def guard_hops(span):
    hops = span.attributes.get("orchestration.hop_count", 0)
    if hops > MAX_HOPS:
        emit_alert("runaway_loop", trace_id=span.context.trace_id, hops=hops)
        raise RunawayLoopError(f"{hops} hops exceeds budget {MAX_HOPS}")
WARNING
Watch Out
A failing agent rarely throws — it returns a confident, wrong answer with a clean trace, so you cannot alert on exceptions alone. Pair runtime metrics with the offline evals from the Agentic AI course and run a sampled eval continuously in production so regressions surface as a metric, not a customer complaint.

Finally, send everything to an OpenTelemetry Collector rather than pointing your app at a backend directly. The collector decouples you from vendor choice, batches and retries exports, and centralises sampling and PII redaction:

# otel-collector.yaml
receivers:
  otlp:
    protocols: { grpc: { endpoint: 0.0.0.0:4317 } }
processors:
  batch: {}
  attributes/redact:
    actions: [ { key: gen_ai.prompt, action: delete } ]  # strip prompts
exporters:
  otlp/traces: { endpoint: tempo:4317, tls: { insecure: true } }
  prometheus:  { endpoint: 0.0.0.0:8889 }
service:
  pipelines:
    traces:  { receivers: [otlp], processors: [batch, attributes/redact], exporters: [otlp/traces] }
    metrics: { receivers: [otlp], processors: [batch], exporters: [prometheus] }

Point Tempo/Jaeger at the traces pipeline for the agent-tree waterfall and Grafana at the Prometheus endpoint for per-agent dashboards. The waterfall is your primary debugging surface — one glance shows fan-out width, the slowest worker, and whether a sub-agent is recursing.

TIP
Tip
Build one dashboard row per agent showing latency p95, token rate, and success rate side by side, plus a fleet-wide trace-depth heatmap. When something breaks at 3am, the misbehaving specialist lights up red in its own row instead of being buried in an aggregate.

Questions & Answers

Q: Won't tracing every agent and tool call add latency and cost to an already-expensive system?
Span creation is cheap and the OTLP exporter batches asynchronously off the request path, so the inline cost is microseconds. The real cost is storage, which you control with sampling — trace 100% of errors and escalations but only a small percentage of successful runs. The token cost of an agent dwarfs the cost of observing it.
Q: How do I keep PII and secrets out of traces when prompts and tool inputs are the most useful thing to log?
Default to off — capture schema, field names, and payload sizes, not values. The Anthropic Agent SDK gates content capture behind flags like OTEL_LOG_TOOL_CONTENT for this reason. Redact at the collector with an attributes processor so it cannot be bypassed by one misconfigured service, and enable full-content capture only for scoped, time-boxed debugging.
Q: My agents run as separate workers behind a queue — how does one trace span across processes?
W3C trace context propagation. The producer calls inject() to serialize the active context onto the message; the worker calls extract() to restore it before opening its span. The Agent SDK does this automatically by reading TRACEPARENT and TRACESTATE from the environment in non-interactive sessions.
Q: Why shouldn't I just alert on the model API returning errors?
Because the dangerous failures return success. A supervisor routing to the wrong specialist, an agent that loops, or a prompt regression all produce clean 200s and tidy traces. Measure task success at the orchestration layer and run sampled evals continuously — a flat API error rate plus a falling task success rate is the signature of silent quality drift.

Key Takeaways

  1. Agents need decision-level observability — capture the prompt, tool inputs, reasoning, and chosen branch, not just timing, because bad decisions look like successful requests.
  2. Use the GenAI semantic conventions — standard span names (invoke_agent, invoke_workflow) and gen_ai.* attributes keep your telemetry portable, but pin the version since they are still in Development.
  3. Propagate one trace across the whole fleet — W3C trace context via inject/extract (and the Agent SDK's TRACEPARENT handling) turns a scattered set of calls into one debuggable tree.
  4. Dimension metrics per agent, control cardinality — latency, tokens, and success rate per agent drive your dashboards; keep high-cardinality labels off metrics and on traces.
  5. Plan for replay and drift, not just exceptions — record replayable state for sampled and failed runs, and alert on runaway loops and falling task success rate, which exceptions will never reveal.

Next Steps: Lesson 9: Cost Management & Optimization