Building Your First Agent

60 min intermediate Lesson 6

Learning Outcomes

  • Build a complete agent loop (think → act → observe → repeat) on Claude's Messages API
  • Define client tools with correct JSON Schema and dispatch their execution in code
  • Handle the four failure modes: tool errors, bad input, infinite loops, and token limits
  • Maintain conversation memory across turns so the agent accumulates state
  • Test the agent on a task ladder and read its trace to debug it

Lesson Plan

Segment Duration Topic
Intro 3 min What we're building and the loop
Setup 7 min Project, SDK, API key, first call
Define tools 10 min Three client tools with real JSON Schema
The loop 14 min Append, dispatch, return tool_result; the list as memory
Edge cases 14 min Tool failures, bad input, loop caps, token budget
Testing 9 min Server tool, task ladder, reading the trace
Wrap-up 3 min Key takeaways, preview safety

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ and the official SDK: pip install anthropic
  • An Anthropic API key in ANTHROPIC_API_KEY (console.anthropic.com)
  • A terminal and a code editor

1 Project Setup and First Call

An agent is a while loop around a model call. First set up the project:

mkdir first-agent && cd first-agent
python -m venv .venv
source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install anthropic
export ANTHROPIC_API_KEY="sk-ant-..."   # Windows: set ANTHROPIC_API_KEY=...

A one-line smoke test confirms the SDK, key, and model are wired up. The client reads ANTHROPIC_API_KEY from the environment, so a bare call needs no config:

import anthropic
client = anthropic.Anthropic()
resp = client.messages.create(model="claude-opus-4-8", max_tokens=64,
    messages=[{"role": "user", "content": "Reply with one word: ready"}])
print(resp.content[0].text, resp.stop_reason)   # -> ready end_turn
NOTE
The mental model
An LLM call is stateless. The agent loop is the only thing that carries state forward, by re-sending the growing message list on every call.
WARNING
Never hardcode keys
The SDK reads ANTHROPIC_API_KEY from the environment. Do not paste your key into source — it ends up in git history.

2 Define the Tools

A tool is a function the model can ask you to run, described with a JSON Schema. Our three client tools run in your process — you get a tool_use request and execute it. Each definition has three fields: name (matching ^[a-zA-Z0-9_-]{1,64}$), a description, and an input_schema.

TOOLS = [
    {
        "name": "calculator",
        "description": "Evaluate an arithmetic expression and return the "
            "result. Use this for ANY math instead of computing it yourself, "
            "since you make mistakes. Supports + - * / ** and parentheses.",
        "input_schema": {
            "type": "object",
            "properties": {
                "expression": {"type": "string",
                               "description": "e.g. '(1234 * 7) / 3'"}
            },
            "required": ["expression"],
        },
    },
    # ... plus read_file and write_file, same shape ...
]

Define read_file (required path) and write_file (required path, content) the same way. The description is the most important field — the only place Claude learns when to reach for a tool. Note how calculator explicitly tells the model to stop doing mental math.

TIP
Descriptions are prompts
Aim for three to four sentences per tool: what it does, when to use it, when NOT to, and what it returns on failure. A vague 'does math' produces unreliable selection.

3 Implement the Tool Dispatcher

The schema tells Claude how to call a tool; you still need code that runs it. Map each name to a function and route through one dispatcher that never raises — a failed tool is normal input, not a crash.

from asteval import Interpreter   # pip install asteval; arithmetic-only evaluator
from pathlib import Path

_calc = Interpreter()   # no imports, no attribute access — safe for model input

def calculator(expression: str) -> str:
    return str(_calc(expression))

# read_file / write_file are thin wrappers over Path(path).read_text / write_text
REGISTRY = {"calculator": calculator,
            "read_file": read_file, "write_file": write_file}

def dispatch(name: str, args: dict) -> tuple[str, bool]:
    """Returns (result_text, is_error). Never raises."""
    fn = REGISTRY.get(name)
    if fn is None:
        return f"Unknown tool: {name}", True
    try:
        return fn(**args), False
    except Exception as e:
        return f"{type(e).__name__}: {e}", True   # instructive error for Claude

The (text, is_error) tuple maps onto the API's tool_result block, built next.

WARNING
Never eval() model output
Running Python's eval() on a model string is remote code execution waiting to happen — it could be steered into __import__('os').system('rm -rf ~'). Use an arithmetic-only evaluator. Sandboxing is Lesson 7.
TIP
Write instructive errors
Return a message that tells Claude what to do next — config.json not found; list the directory first beats a bare failed. Claude reads it and self-corrects.

4 Build the Agent Loop

The loop runs think → act → observe → repeat. Critically, append Claude's entire assistant message (text and tool_use blocks), then reply with a user message whose tool_result blocks come first, matched by tool_use_id:

def run_agent(goal: str, max_steps: int = 12) -> str:
    messages = [{"role": "user", "content": goal}]
    for step in range(max_steps):
        resp = client.messages.create(
            model="claude-opus-4-8", max_tokens=2048,
            tools=TOOLS, messages=messages)

        messages.append({"role": "assistant", "content": resp.content})  # persist
        if resp.stop_reason != "tool_use":             # no tools -> final answer
            return "".join(b.text for b in resp.content if b.type == "text")

        results = []                                   # one tool_result per call
        for block in resp.content:
            if block.type == "tool_use":
                text, is_error = dispatch(block.name, block.input)
                print(f"  [step {step}] {block.name}({block.input}) -> {text[:60]}")
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": text, "is_error": is_error})

        messages.append({"role": "user", "content": results})  # tool_result FIRST
    return "Agent hit the step limit without finishing."

That is a complete agent. The tool_use block carries id, name, input; the matching tool_result carries tool_use_id, content, is_error. Drive it: run_agent("Read prices.txt, sum each line, write the total to out.txt").

A single assistant turn can contain several tool_use blocks (parallel tool use); the loop handles that, returning one tool_result per call.

WARNING
Order is enforced
Every tool_result must come before any text in the result user message, with no message between the tool_use and your tool_result. Violate this and the API returns a 400.

The messages list is the agent's working memory — no hidden store. Every cycle appends to it and the whole list is re-sent, which is why the agent "remembers" it read the file two steps ago. That is short-term memory. For long-term memory across restarts (Lesson 4: Memory Systems), serialize the list to JSON (call .model_dump() on each resp.content block, since they are SDK objects) and reload.

TIP
System prompt vs memory
Persona and standing rules belong in the system parameter of messages.create, not the message list. The prompt steers; the list remembers.

5 Handle the Four Edge Cases

A demo agent works once; a real agent survives failure. Four things go wrong, and three are already handled: a tool that throws returns is_error=True from the dispatcher (Step 3); bad input comes back as an error string Claude retries against; an infinite loop is bounded by max_steps (Step 4). The fourth — token blow-up — needs management: each loop re-sends the whole history, so a long run can exceed the context window. Read usage each iteration and, once it crosses a budget, compact:

if resp.usage.input_tokens + resp.usage.output_tokens > 150_000:
    messages = compact(messages)   # summarize old turns, keep recent ones

compact() is another model call: summarize messages [1:-2] into a paragraph, then rebuild as [goal, summary, recent_turns] — like /compact.

A subtler loop is the non-terminating retry: a tool keeps failing and Claude keeps re-calling it with the same bad input. Track recent (tool_name, sorted-json input) signatures; if one repeats three times, inject a tool_result telling the agent to stop and ask the user. These limits are the simplest agent safety — a resource cap Lesson 7 turns into a full guardrail layer.

WARNING
Always cap the loop
Never run an agent loop without a hard max_steps bound. A model that misreads a tool error can loop until your token budget is gone — a real, expensive incident.

6 Add a Server Tool and Test on a Ladder

You can mix in server tools — tools Anthropic runs on its own infrastructure. Add them by type; Claude executes them with no tool_result handling from you, returning results in the same turn. Web search is the canonical one, with a built-in max_uses cap:

TOOLS = TOOLS + [
    {"type": "web_search_20250305", "name": "web_search", "max_uses": 5}
]

Now test on a task ladder of harder goals:

  1. "What is 17 * 6 + 4?" — one tool call, then end_turn.
  2. "Sum the numbers in data.txt, write the total to out.txt" — chained tools.
  3. "Find the latest Python release; save the version to py.txt" — server plus client tool.
  4. "Read a file that doesn't exist, then recover" — error handling, self-correction.

Run each rung and read the trace — the [step N] tool(args) -> result lines are your debugger. Most bugs are visible there: a wrong argument, an ignored result, or a non-terminating loop.

TIP
The trace is the truth
When an agent does something baffling, never guess — read the step trace. Its behavior is always explained by what it saw in the tool results. Lesson 8 turns this into metrics like step-efficiency and completion rate.

Questions & Answers

Q: My agent stops after one tool call instead of continuing. Why?
Almost always you forgot to append the assistant message before the tool_result, or your tool_use_id doesn't match the original tool_use.id. Push resp.content onto messages, then a user message with the matching tool_result. A mismatch ends the chat or 400s.
Q: How do I stop a runaway agent from burning my whole token budget?
Two hard limits, always: a max_steps cap and a token budget checked against resp.usage each iteration, plus a stuck-detector for repeated identical calls. These resource guardrails are a lightweight version of Lesson 7.
Q: Is returning is_error: true better than raising an exception?
Yes, for tool failures. A raised exception kills your loop; an is_error result is information the agent can act on — Claude reads it and retries before giving up. Reserve exceptions for harness bugs, not expected failures like a missing file.
Q: Do I have to write this loop by hand every time?
No — frameworks (and Anthropic's own Tool Runner) automate the append/dispatch/return cycle. But build it by hand once: when a production agent misbehaves you need to understand the loop to debug it, and frameworks hide the very mechanics you'll need.

Key Takeaways

  1. An agent is a loop, not an object. Think → act → observe → repeat — the while loop around messages.create is the architecture.
  2. The messages list is the memory. Calls are stateless; state persists only because you append every turn and re-send the history. Serialize it for cross-session memory.
  3. Schemas teach selection; dispatchers run code. The description is the highest-leverage thing you write. Route names through one dispatcher returning (text, is_error) that never raises.
  4. Match every tool_use_id, put tool_result first. The most common bug is a broken append/return cycle: persist the assistant turn, then reply with tool_result first.
  5. Cap everything. A max_steps bound, a token budget, max_uses on server tools, and a stuck-detector separate an agent from a costly incident.
  6. Test on a ladder, read the trace. Debug by reading what the agent observed, not by guessing.

Next Steps: Lesson 7: Agent Safety & Guardrails