Building Your First Agent
Learning Outcomes
- Build a complete agent loop (think → act → observe → repeat) on Claude's Messages API
- Define client tools with correct JSON Schema and dispatch their execution in code
- Handle the four failure modes: tool errors, bad input, infinite loops, and token limits
- Maintain conversation memory across turns so the agent accumulates state
- Test the agent on a task ladder and read its trace to debug it
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | What we're building and the loop |
| Setup | 7 min | Project, SDK, API key, first call |
| Define tools | 10 min | Three client tools with real JSON Schema |
| The loop | 14 min | Append, dispatch, return tool_result; the list as memory |
| Edge cases | 14 min | Tool failures, bad input, loop caps, token budget |
| Testing | 9 min | Server tool, task ladder, reading the trace |
| Wrap-up | 3 min | Key takeaways, preview safety |
Before You Begin
Pre-work:
- Complete Lesson 3: Tool Use — Giving Agents Hands for the tool-calling mechanics
- Skim Lesson 5: Planning & Reasoning — the agent plans inside the loop
- Be comfortable with Python 3.10+ and JSON
Shopping List:
- Python 3.10+ and the official SDK:
pip install anthropic - An Anthropic API key in
ANTHROPIC_API_KEY(console.anthropic.com) - A terminal and a code editor
An agent is a while loop around a model call. First set up the project:
mkdir first-agent && cd first-agent
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install anthropic
export ANTHROPIC_API_KEY="sk-ant-..." # Windows: set ANTHROPIC_API_KEY=...
A one-line smoke test confirms the SDK, key, and model are wired up. The client reads ANTHROPIC_API_KEY from the environment, so a bare call needs no config:
import anthropic
client = anthropic.Anthropic()
resp = client.messages.create(model="claude-opus-4-8", max_tokens=64,
messages=[{"role": "user", "content": "Reply with one word: ready"}])
print(resp.content[0].text, resp.stop_reason) # -> ready end_turn
ANTHROPIC_API_KEY from the environment. Do not paste your key into source — it ends up in git history.A tool is a function the model can ask you to run, described with a JSON Schema. Our three client tools run in your process — you get a tool_use request and execute it. Each definition has three fields: name (matching ^[a-zA-Z0-9_-]{1,64}$), a description, and an input_schema.
TOOLS = [
{
"name": "calculator",
"description": "Evaluate an arithmetic expression and return the "
"result. Use this for ANY math instead of computing it yourself, "
"since you make mistakes. Supports + - * / ** and parentheses.",
"input_schema": {
"type": "object",
"properties": {
"expression": {"type": "string",
"description": "e.g. '(1234 * 7) / 3'"}
},
"required": ["expression"],
},
},
# ... plus read_file and write_file, same shape ...
]
Define read_file (required path) and write_file (required path, content) the same way. The description is the most important field — the only place Claude learns when to reach for a tool. Note how calculator explicitly tells the model to stop doing mental math.
The schema tells Claude how to call a tool; you still need code that runs it. Map each name to a function and route through one dispatcher that never raises — a failed tool is normal input, not a crash.
from asteval import Interpreter # pip install asteval; arithmetic-only evaluator
from pathlib import Path
_calc = Interpreter() # no imports, no attribute access — safe for model input
def calculator(expression: str) -> str:
return str(_calc(expression))
# read_file / write_file are thin wrappers over Path(path).read_text / write_text
REGISTRY = {"calculator": calculator,
"read_file": read_file, "write_file": write_file}
def dispatch(name: str, args: dict) -> tuple[str, bool]:
"""Returns (result_text, is_error). Never raises."""
fn = REGISTRY.get(name)
if fn is None:
return f"Unknown tool: {name}", True
try:
return fn(**args), False
except Exception as e:
return f"{type(e).__name__}: {e}", True # instructive error for Claude
The (text, is_error) tuple maps onto the API's tool_result block, built next.
eval() on a model string is remote code execution waiting to happen — it could be steered into __import__('os').system('rm -rf ~'). Use an arithmetic-only evaluator. Sandboxing is Lesson 7.config.json not found; list the directory first beats a bare failed. Claude reads it and self-corrects.The loop runs think → act → observe → repeat. Critically, append Claude's entire assistant message (text and tool_use blocks), then reply with a user message whose tool_result blocks come first, matched by tool_use_id:
def run_agent(goal: str, max_steps: int = 12) -> str:
messages = [{"role": "user", "content": goal}]
for step in range(max_steps):
resp = client.messages.create(
model="claude-opus-4-8", max_tokens=2048,
tools=TOOLS, messages=messages)
messages.append({"role": "assistant", "content": resp.content}) # persist
if resp.stop_reason != "tool_use": # no tools -> final answer
return "".join(b.text for b in resp.content if b.type == "text")
results = [] # one tool_result per call
for block in resp.content:
if block.type == "tool_use":
text, is_error = dispatch(block.name, block.input)
print(f" [step {step}] {block.name}({block.input}) -> {text[:60]}")
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": text, "is_error": is_error})
messages.append({"role": "user", "content": results}) # tool_result FIRST
return "Agent hit the step limit without finishing."
That is a complete agent. The tool_use block carries id, name, input; the matching tool_result carries tool_use_id, content, is_error. Drive it: run_agent("Read prices.txt, sum each line, write the total to out.txt").
A single assistant turn can contain several tool_use blocks (parallel tool use); the loop handles that, returning one tool_result per call.
tool_result must come before any text in the result user message, with no message between the tool_use and your tool_result. Violate this and the API returns a 400.The messages list is the agent's working memory — no hidden store. Every cycle appends to it and the whole list is re-sent, which is why the agent "remembers" it read the file two steps ago. That is short-term memory. For long-term memory across restarts (Lesson 4: Memory Systems), serialize the list to JSON (call .model_dump() on each resp.content block, since they are SDK objects) and reload.
system parameter of messages.create, not the message list. The prompt steers; the list remembers.A demo agent works once; a real agent survives failure. Four things go wrong, and three are already handled: a tool that throws returns is_error=True from the dispatcher (Step 3); bad input comes back as an error string Claude retries against; an infinite loop is bounded by max_steps (Step 4). The fourth — token blow-up — needs management: each loop re-sends the whole history, so a long run can exceed the context window. Read usage each iteration and, once it crosses a budget, compact:
if resp.usage.input_tokens + resp.usage.output_tokens > 150_000:
messages = compact(messages) # summarize old turns, keep recent ones
compact() is another model call: summarize messages [1:-2] into a paragraph, then rebuild as [goal, summary, recent_turns] — like /compact.
A subtler loop is the non-terminating retry: a tool keeps failing and Claude keeps re-calling it with the same bad input. Track recent (tool_name, sorted-json input) signatures; if one repeats three times, inject a tool_result telling the agent to stop and ask the user. These limits are the simplest agent safety — a resource cap Lesson 7 turns into a full guardrail layer.
max_steps bound. A model that misreads a tool error can loop until your token budget is gone — a real, expensive incident.You can mix in server tools — tools Anthropic runs on its own infrastructure. Add them by type; Claude executes them with no tool_result handling from you, returning results in the same turn. Web search is the canonical one, with a built-in max_uses cap:
TOOLS = TOOLS + [
{"type": "web_search_20250305", "name": "web_search", "max_uses": 5}
]
Now test on a task ladder of harder goals:
- "What is 17 * 6 + 4?" — one tool call, then
end_turn. - "Sum the numbers in data.txt, write the total to out.txt" — chained tools.
- "Find the latest Python release; save the version to py.txt" — server plus client tool.
- "Read a file that doesn't exist, then recover" — error handling, self-correction.
Run each rung and read the trace — the [step N] tool(args) -> result lines are your debugger. Most bugs are visible there: a wrong argument, an ignored result, or a non-terminating loop.
Questions & Answers
tool_use_id doesn't match the original tool_use.id. Push resp.content onto messages, then a user message with the matching tool_result. A mismatch ends the chat or 400s.max_steps cap and a token budget checked against resp.usage each iteration, plus a stuck-detector for repeated identical calls. These resource guardrails are a lightweight version of Lesson 7.is_error: true better than raising an exception?is_error result is information the agent can act on — Claude reads it and retries before giving up. Reserve exceptions for harness bugs, not expected failures like a missing file.Key Takeaways
- An agent is a loop, not an object. Think → act → observe → repeat — the
whileloop aroundmessages.createis the architecture. - The messages list is the memory. Calls are stateless; state persists only because you append every turn and re-send the history. Serialize it for cross-session memory.
- Schemas teach selection; dispatchers run code. The
descriptionis the highest-leverage thing you write. Route names through one dispatcher returning(text, is_error)that never raises. - Match every
tool_use_id, puttool_resultfirst. The most common bug is a broken append/return cycle: persist the assistant turn, then reply withtool_resultfirst. - Cap everything. A
max_stepsbound, a token budget,max_useson server tools, and a stuck-detector separate an agent from a costly incident. - Test on a ladder, read the trace. Debug by reading what the agent observed, not by guessing.
Next Steps: Lesson 7: Agent Safety & Guardrails