The Anthropic Agent SDK

50 min intermediate Lesson 5

Learning Outcomes

  • Install the Claude Agent SDK and run an agent loop in Python
  • Define specialist subagents with scoped tools using AgentDefinition
  • Wire a coordinator to delegate via Agent and track handoffs by parent_tool_use_id
  • Share context through sessions and an in-process MCP blackboard
  • Enforce orchestration-layer guardrails with permissions and lifecycle hooks

Lesson Plan

Segment Duration Topic
Intro 3 min Where the SDK fits in your orchestration stack
Setup 9 min Installing the SDK and the agent loop
Build 12 min Subagents and the coordinator handoff
Build 12 min Shared context and orchestration guardrails
Debug 9 min Tracing runs and common failure modes
Wrap-up 5 min Trade-offs and what comes next

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ and an Anthropic API key exported as ANTHROPIC_API_KEY
  • A terminal, a scratch project directory, and a small API budget for live runs

1 Install the SDK and Run the Agent Loop

The Claude Agent SDK is the same agent loop, tool execution, and context management that power Claude Code, as a Python/TypeScript library with built-in tools (Read, Edit, Bash, Glob, Grep, WebSearch, WebFetch). Unlike the lower-level Client SDK, it runs the tool loop for you.

pip install claude-agent-sdk   # requires Python 3.10+
export ANTHROPIC_API_KEY=your-api-key

The core primitive is query() — an async generator yielding each message (assistant turns, tool calls, results):

import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions

async def main():
    async for message in query(prompt="What files are here?",
            options=ClaudeAgentOptions(allowed_tools=["Bash", "Glob"])):
        if hasattr(message, "result"):
            print(message.result)

asyncio.run(main())

allowed_tools pre-approves tools so the loop runs unattended; anything else triggers a permission decision (Step 5). For multi-turn sessions, use ClaudeSDKClient.

NOTE
Where this sits in your stack
The SDK runs the loop inside your process, session state as JSONL on disk. Anthropic also offers Managed Agents — a hosted REST API that runs the loop and sandbox for you. Prototype locally with the SDK, then move to Managed Agents for production.

2 Define Specialist Subagents

We are building a research workflow: a coordinator, search, synthesis, and fact-checker agent. The SDK's unit of specialization is the subagent, declared with AgentDefinition — each gets isolated context and scoped tools, the decomposition discipline from Lesson 4.

from claude_agent_sdk import AgentDefinition

AGENTS = {
    "searcher": AgentDefinition(
        description="Finds primary sources for a focused question.",
        prompt="Search the web and return JSON url/claim/quote objects. Do not synthesize.",
        tools=["WebSearch", "WebFetch"]),
    "synthesizer": AgentDefinition(
        description="Combines sources into a structured briefing.",
        prompt="Group claims by theme, cite each by url, add no facts absent from sources.",
        tools=["Read"]),
    "fact_checker": AgentDefinition(
        description="Independently verifies each claim against its source.",
        prompt="Fetch each claim's url; mark SUPPORTED/UNSUPPORTED/UNCERTAIN. Return the table.",
        tools=["WebFetch"]),
}

Each tools list is a strict subset of system capabilities — least-privilege at the agent boundary, and the topology is legible straight off the lists. Context is isolated per subagent: it sees only what is passed in and returns only what is relevant, so pass inputs in.

TIP
Keep outputs structured
A subagent returning free prose forces the next one to re-parse it — an unreliable, token-heavy handoff. Specify an output contract (JSON shape, verdict table). Structured artifacts are how Anthropic's production long-running-agent pattern hands off between its Planner, Generator, and Evaluator stages.

3 The Coordinator and the Handoff

Subagents are invoked through the built-in Agent tool. Include Agent in allowed_tools and register specialists under agents so the coordinator delegates without approving every hop. This is the supervisor pattern from Lesson 2: one coordinator, many workers.

import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions

COORDINATOR = "Decompose; delegate to searcher/synthesizer/fact_checker; drop UNSUPPORTED claims; never do the work."

async def run_research(question: str):
    opts = ClaudeAgentOptions(
        system_prompt=COORDINATOR,
        allowed_tools=["Agent"],   # its only tool: delegate
        agents=AGENTS, permission_mode="default")
    async for message in query(prompt=question, options=opts):
        tag = getattr(message, "parent_tool_use_id", None)
        prefix = f"[subagent {tag}]" if tag else "[coordinator]"
        if hasattr(message, "result"):
            print(prefix, message.result)

asyncio.run(run_research("State of solid-state battery commercialization?"))

The coordinator's only move is delegation. Each message inside a subagent's run carries a parent_tool_use_id, attributing every line to the execution that produced it — the backbone of tracing (Step 6, Lesson 8). Branches are fault-isolated: a failed search fails one sub-question.

NOTE
Delegation is an async fan-out point
When the coordinator delegates independent sub-questions, those runs proceed in parallel — the SDK supports multiple subagents at once. Treat each branch like a queue job: independent, retryable, reassembled by the supervisor.

4 Shared Context: Sessions and a Blackboard

Two mechanisms share state. Sessions give continuity: capture the session ID from the init message, then resume or fork via resume=session_id. The second is the blackboard from Lesson 3: agents read and write a common store instead of passing everything through the prompt. Expose it as an in-process MCP tool via tool and create_sdk_mcp_server:

from claude_agent_sdk import tool, create_sdk_mcp_server, ClaudeAgentOptions

ARTIFACTS = {}  # in production: Redis, S3, or a database

@tool("save_artifact", "Persist a named JSON artifact", {"key": str, "value": str})
async def save(args):
    ARTIFACTS[args["key"]] = args["value"]
    return {"content": [{"type": "text", "text": f"saved {args['key']}"}]}

@tool("load_artifact", "Read an artifact by key", {"key": str})
async def load(args):  # MISSING sentinel lets readers degrade gracefully
    return {"content": [{"type": "text", "text": ARTIFACTS.get(args["key"], "MISSING")}]}

store = create_sdk_mcp_server("blackboard", "1.0.0", tools=[save, load])
options = ClaudeAgentOptions(mcp_servers={"blackboard": store},
    allowed_tools=["mcp__blackboard__save_artifact", "mcp__blackboard__load_artifact"])

The searcher writes sources:battery; the synthesizer reads it — by reference, not a blob in context.

WARNING
The blackboard is eventually consistent
A plain dict (or even Redis) gives no transaction across a run. If two subagents write the same key, last-writer-wins. Namespace keys per run and never assume a parallel branch's write has landed before you read.

5 Guardrails at the Orchestration Layer

Course 04 covered per-agent safety. Orchestration adds a tier the coordinator enforces on its workers. Permissions are coarse: allowed_tools pre-approves the safe set, permission_mode decides the rest. Hooks are fine-grained — callbacks at lifecycle points (PreToolUse, PostToolUse, Stop, and more) to validate, log, block, or transform behavior. A PreToolUse hook is your enforcement point: inspect the input and refuse before it runs.

from claude_agent_sdk import ClaudeAgentOptions, HookMatcher

ALLOWED = {"docs.anthropic.com", "arxiv.org", "nature.com"}

async def restrict_fetch(input_data, tool_use_id, context):
    url = input_data.get("tool_input", {}).get("url", "")
    if url and not any(h in url for h in ALLOWED):   # a deny blocks the call
        return {"hookSpecificOutput": {"permissionDecision": "deny"}}
    return {}

opts = ClaudeAgentOptions(allowed_tools=["Agent", "WebFetch"], agents=AGENTS,
    hooks={"PreToolUse": [HookMatcher(matcher="WebFetch", hooks=[restrict_fetch])]})

Hooks run in-process and fire for subagent calls too, so one allowlist enforces a network policy across the fleet; a PostToolUse audit hook adds the trail.

TIP
Deterministic guardrails over instructions
A system-prompt line is a request; a PreToolUse hook is enforcement. Anything with blast radius — spend, destructive Bash — belongs in a hook or permission, not an instruction.

6 Tracing Runs and Common Failure Modes

A multi-agent run is interleaved — branches stream concurrently and a flat log is unreadable. Bucket each message by parent_tool_use_id (defaulting to "coordinator") for one timeline per subagent. Common failures and their signatures:

Symptom Likely cause First move
Coordinator answers itself, never delegates Agent missing from allowed_tools Strip it to Agent only
Subagent stalls or loops Needed tool absent from its AgentDefinition.tools Check the per-agent list, not the global
Synthesis invents facts Sources passed as prose, or a branch dropped Enforce JSON; verify branches landed
Branch silently empty Subagent errored; isolated context hid it Add a Stop hook to surface errors

Log the artifact at each handoff boundary keyed by parent_tool_use_id — these become OpenTelemetry spans in Lesson 8.


Questions & Answers

Q: We run Temporal/Prefect. Why the SDK, not just calling the API from activities?
Different layers. The engine owns durability and deterministic flow; the SDK owns the non-deterministic inner loop, so you don't reimplement it per activity. The pattern wraps an SDK query() inside a workflow activity with retries — what Lesson 6: Workflow Engines builds.
Q: Subagents have isolated context. How do I get idempotency for retries?
You won't get token-level determinism. Make the effect idempotent: route writes through the blackboard with run-namespaced keys so a retry overwrites its own slot, and make reassembly a pure function of the artifacts present.
Q: If a subagent silently fails, does the run hang or produce garbage?
A failed branch can return an empty result the coordinator may not notice — isolated context hides it. Add a Stop/PostToolUse hook to surface errors, assert each artifact exists before synthesis, and define a fallback (Lesson 7).
Q: Can I trust a hook to actually stop a dangerous action, or is it advisory?
A PreToolUse deny decision blocks the call before it executes — enforcement, not a suggestion. Guardrails with real blast radius (network allowlists, destructive Bash, spend) belong in hooks, not prompt text.

Key Takeaways

  1. The SDK runs the loop so you compose agents. Built-in tools, context management, and delegation come free — you orchestrate instead of re-implementing the tool-call loop.
  2. Subagents are your decomposition unit. Each AgentDefinition has isolated context and scoped tools; the coordinator delegates through Agent — the supervisor pattern with least privilege.
  3. parent_tool_use_id is the spine of observability. Every subagent message is attributable to the delegation that produced it — group on it for traces and debugging.
  4. Share context deliberately. Sessions give continuity; an MCP blackboard gives handoff-by-reference. Both are eventually consistent — namespace per run.
  5. Guard at the orchestration layer, and design for partial failure. Permissions and hooks impose deterministic policy across the fleet; assert artifacts exist before reassembly.

Next Steps: Lesson 6: Workflow Engines