Frameworks & SDK Reference
Two layers combine in production multi-agent systems: the agent runtime (the Anthropic Agent SDK, which runs Claude's tool loop inside your process) and the durable workflow engine (Temporal, Prefect, Airflow, which keeps a long-running plan alive across crashes, retries, and restarts). They solve orthogonal problems — the SDK gives an agent autonomy; the engine gives the surrounding orchestration durability, scheduling, and exactly-once-style semantics.
| Concern | Agent SDK | Workflow engine |
|---|---|---|
| Non-deterministic reasoning + tool loop | Yes (this is the agent) | No (this is what you wrap) |
| Durable state across process restarts | No — JSONL on local disk | Yes — persisted run history |
| Cross-process retries / backoff | No (per-tool only) | Yes — first-class, per step |
| Long-running / multi-day workflows | Not designed for it | Yes — seconds to years |
| Human-in-the-loop pauses | In-run only (AskUserQuestion) |
Durable signals survive restarts |
| Fan-out / parallel branches | Subagents in one process | Activities across a worker pool |
| Scheduling (cron, event triggers) | No | Yes |
Anthropic Agent SDK
The Agent SDK packages the same agent loop, built-in tools, and context management that power Claude Code as a library, in Python and TypeScript. It runs in your own process and infrastructure; session state is JSONL on your filesystem.
| Abstraction | Python API | Purpose |
|---|---|---|
| One-shot query | query(prompt, options) |
Async iterator over response messages |
| Interactive client | ClaudeSDKClient |
Bidirectional sessions; enables custom tools + hooks |
| Configuration | ClaudeAgentOptions |
allowed_tools, system_prompt, max_turns, cwd, permission_mode, resume, mcp_servers, hooks, agents |
| Custom tool | @tool(name, desc, schema) |
In-process MCP tool defined as a Python function |
| Subagent | AgentDefinition |
Specialized agent invoked via the Agent tool |
| Lifecycle hooks | HookMatcher |
PreToolUse, PostToolUse, Stop, SessionStart, SessionEnd, UserPromptSubmit |
Install with pip install claude-agent-sdk (Python 3.10+). Authenticate with ANTHROPIC_API_KEY, or route through Amazon Bedrock, Google Vertex AI, or Azure via environment flags.
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions, AgentDefinition
async def main():
async for message in query(
prompt="Use the code-reviewer agent to review this codebase",
options=ClaudeAgentOptions(
allowed_tools=["Read", "Glob", "Grep", "Agent"],
agents={
"code-reviewer": AgentDefinition(
description="Expert code reviewer for quality and security.",
prompt="Analyze code quality and suggest improvements.",
tools=["Read", "Glob", "Grep"],
)
},
),
):
if hasattr(message, "result"):
print(message.result)
asyncio.run(main())
Subagents are the SDK's native fan-out primitive: the main agent delegates focused subtasks and they report back. Messages from a subagent carry a parent_tool_use_id, so you can attribute tokens and trace each branch. For fully hosted execution without operating your own sandbox or session store, Anthropic also offers Managed Agents (a hosted REST API where Anthropic runs the loop and a per-session sandbox); a common path is to prototype with the SDK locally, then move to Managed Agents for production.
What the SDK does NOT give you: durable orchestration. A crash mid-run loses in-flight work beyond the last JSONL checkpoint; there is no cross-process retry, no scheduler, and no exactly-once guarantee around the whole plan. That gap is exactly what a workflow engine fills.
Workflow engine comparison
| Engine | What it is for | Programming model | Durability mechanism | Strengths | Combine with agents when |
|---|---|---|---|---|---|
| Temporal | Durable execution of arbitrary long-running business logic | Imperative code (Workflows + Activities) in Python/Go/TS/Java | Event-sourced history; deterministic workflow replay | Strongest fault tolerance; multi-day waits; durable signals; per-Activity retries | You need correctness guarantees, human-in-loop pauses, or workflows spanning hours to days |
| Prefect | Resilient Python data/ML pipelines | Decorated @flow / @task Python functions |
Run-state tracking + result persistence | Pythonic, low ceremony; caching, retries, event-driven automations | The orchestration is mostly dynamic Python and the team wants minimal infra |
| Airflow | Scheduled batch DAGs / data engineering | DAGs of operators; scheduler triggers tasks | Metadata DB of task instances + retries | Mature ecosystem; rich scheduling (cron, timetables, assets); operators | Agent steps slot into an existing batch/ETL DAG on a schedule |
Choosing between them
- Temporal — orchestration must be bulletproof: financial flows, transactions with compensation, agents that pause for human approval for hours, any plan where losing state is unacceptable. A Temporal Workflow Execution runs effectively once and to completion whether that takes seconds or years.
- Prefect — orchestration is dynamic Python that branches on data, you want decorators not a DSL, and lighter durability than Temporal's replay model is acceptable.
- Airflow — agent work is one stage in a scheduled, mostly-deterministic data pipeline and you already run a scheduler and metadata DB.
Wrapping agent calls as workflow activities
The core pattern is identical across engines: all non-determinism lives in an activity/task, never in the workflow body. Agent calls are non-deterministic (the LLM, tool I/O, network), so they execute outside the engine's deterministic control path. Temporal makes this explicit — workflow code is replayed to reconstruct state, so it must be deterministic; LLM invocations, API calls, and DB queries belong in Activities, which run outside the replay path and are automatically retried.
from datetime import timedelta
from temporalio import activity, workflow
from temporalio.common import RetryPolicy
@activity.defn
async def run_research_agent(topic: str) -> str:
# Non-deterministic agent loop isolated here, outside replay.
from claude_agent_sdk import query, ClaudeAgentOptions
result = ""
async for message in query(
prompt=f"Research {topic} and return a sourced summary.",
options=ClaudeAgentOptions(allowed_tools=["WebSearch", "WebFetch"]),
):
if hasattr(message, "result"):
result = message.result
return result
@workflow.defn
class ResearchWorkflow:
@workflow.run
async def run(self, topic: str) -> str:
# Deterministic orchestration: timeout + bounded retries around the agent.
return await workflow.execute_activity(
run_research_agent, topic,
start_to_close_timeout=timedelta(minutes=10),
retry_policy=RetryPolicy(maximum_attempts=3, backoff_coefficient=2.0),
)
The Prefect equivalent is a decorated task — @task(retries=3, retry_delay_seconds=[5,15,45], timeout_seconds=600) wrapping the same agent loop, fanned out from a @flow with task.submit(...) per item.
Production concerns when bridging the layers
| Concern | Practice |
|---|---|
| Idempotency | Agent activities have side effects (writes, emails, API calls). Pass an idempotency key from the workflow so a retried activity does not double-act. |
| Timeouts | Bound every agent step (start_to_close_timeout / timeout_seconds). Open-ended loops are the top cause of stuck runs. |
| Determinism boundary | Never call the agent, random, uuid4, or the clock from workflow code; use workflow.random() / workflow.uuid4(). |
| Cost budgets | Cap each activity with a token/turn budget (max_turns) so a retry storm cannot run up unbounded spend. |
| Tracing | Propagate the run ID into the agent via system_prompt or hook context; emit OpenTelemetry spans from PostToolUse hooks to correlate tool calls with steps. |
| Human-in-loop | Keep long approval waits in the workflow (durable signals), not the agent process — in-run AskUserQuestion does not survive a crash. |
| Partial failure | Split into multiple activities and persist intermediate outputs, so a failed synthesis step retries without re-running an expensive search step. |
Decision summary
Single autonomous loop, no durability needed -> Agent SDK alone
Durability / retries / multi-day / HITL -> Workflow engine wraps Agent SDK
bulletproof correctness, long pauses -> Temporal
dynamic Python pipelines, low infra -> Prefect
scheduled batch / existing data DAGs -> Airflow
No desire to run sandboxes or session stores -> Managed Agents (hosted)