Tool Use — Giving Agents Hands

45 min intermediate Lesson 3

Learning Outcomes

  • Define tools with correct JSON schemas that models can select and parameterize reliably
  • Implement the tool-use loop: tool_use stop reason, execution, and tool_result return
  • Connect agents to REST APIs, databases, and file systems behind well-designed tools
  • Use the Model Context Protocol (MCP) to expose tools through a standard interface
  • Handle tool failures gracefully with instructive error messages and validation

Lesson Plan

Segment Duration Topic
Intro 3 min Why tools turn LLMs into agents
Concept 6 min The function-calling contract
Build 8 min Defining a tool schema
Build 8 min The agent execution loop
Build 7 min Wrapping APIs, databases, and files
Concept 6 min MCP — the standard interface
Hardening 5 min Error handling and tool design
Wrap-up 2 min Key takeaways, next lesson

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ and the official SDK: pip install anthropic
  • A terminal with node/npx available (for the MCP demo)
  • Optional: a SQLite or PostgreSQL database to wrap as a tool

1 The Function-Calling Contract

A bare LLM can only emit text. Tool use (also called function calling) lets a model ask your code to run a function, then read the result back. The model never executes anything itself — it produces a structured request, your code runs it, and you feed the output back into the conversation.

The Anthropic Messages API splits tools into client tools that run in your application, and server tools (like web_search and code_execution) that run on Anthropic's infrastructure and return results directly. This lesson focuses on client tools — that is where you give an agent hands on your systems.

The full round trip has four moves:

1. You send: messages + a `tools` list (each tool = name + description + input_schema)
2. Claude replies with stop_reason "tool_use" and a tool_use block (id, name, input)
3. You execute the function and send back a tool_result block (matched by tool_use_id)
4. Claude reads the result and either calls another tool or answers

Steps 2 through 4 repeat until the model stops requesting tools — that loop is the ReAct architecture from Lesson 2, made concrete.

NOTE
Key Insight
The model is a planner that emits requests, not an executor. Every side effect happens in code you control. That separation is what makes guardrails (Lesson 7) possible at all.

2 Defining a Tool Schema

A client tool is defined by three fields: a name (matching ^[a-zA-Z0-9_-]{1,64}$), a plaintext description, and an input_schema that is a standard JSON Schema object. The canonical shape:

{
  "name": "get_weather",
  "description": "Get the current weather in a given location. Use this when the user asks about current conditions or temperature for a specific city. Returns temperature and a short summary. Does not return forecasts.",
  "input_schema": {
    "type": "object",
    "properties": {
      "location": {
        "type": "string",
        "description": "The city and state, e.g. San Francisco, CA"
      },
      "unit": {
        "type": "string",
        "enum": ["celsius", "fahrenheit"],
        "description": "Temperature unit; defaults to celsius"
      }
    },
    "required": ["location"]
  }
}

The description is the highest-leverage thing you write: it is the model's only signal for deciding whether to call the tool and how to fill the parameters. Aim for several sentences covering what it does, when to use it (and when not to), and what each parameter means.

Field Required Purpose
name Yes Stable identifier the model emits to call the tool
description Yes The model's only guide to selection and arguments
input_schema Yes JSON Schema; constrains and documents inputs
strict No Set true to guarantee inputs match the schema exactly
TIP
Use the schema as documentation
Add a description to every property and an enum wherever values are constrained. The model treats the schema as a spec — a vague schema produces vague tool calls.
WARNING
Brevity costs you
A description like Gets the price for a ticker leaves the model guessing about format, scope, and timing. Under-specified tools are the number-one cause of wrong or skipped tool calls.

3 The Agent Execution Loop

Now wire the contract into a loop. When Claude wants a tool, the response's stop_reason is "tool_use" and it contains one or more tool_use blocks, each with an id, name, and input. You run the matching function and reply with a user message whose content is a tool_result block keyed by tool_use_id.

import anthropic

client = anthropic.Anthropic()

TOOLS = [ { "name": "get_weather", "description": "...", "input_schema": {...} } ]

def execute_tool(name, tool_input):
    if name == "get_weather":
        # call your real weather code here
        return f"18 degrees, partly cloudy in {tool_input['location']}"
    raise ValueError(f"Unknown tool: {name}")

messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]

while True:
    resp = client.messages.create(
        model="claude-opus-4-8",
        max_tokens=1024,
        tools=TOOLS,
        messages=messages,
    )
    messages.append({"role": "assistant", "content": resp.content})

    if resp.stop_reason != "tool_use":
        break  # Claude produced a final answer

    results = []
    for block in resp.content:
        if block.type == "tool_use":
            output = execute_tool(block.name, block.input)
            results.append({
                "type": "tool_result",
                "tool_use_id": block.id,
                "content": output,
            })
    messages.append({"role": "user", "content": results})

print(resp.content[-1].text)

Two formatting rules the API enforces: tool_result blocks must come first in the user message's content array, and they must immediately follow the assistant turn that requested them — no messages in between.

NOTE
Parallel tool calls
A single assistant turn can contain multiple tool_use blocks. The loop handles this by iterating over every block and returning one tool_result per call before the next round trip.
WARNING
Always bound the loop
Add a max-iteration counter to the while loop. A misbehaving tool or an ambiguous goal can otherwise spin forever. A simple cap belongs here from day one.

4 Wrapping APIs, Databases, and Files

Tools are just functions, so anything callable from your runtime can become one. The art is in the boundary you expose.

A REST API tool — wrap the call, not the raw endpoint:

import requests

def execute_tool(name, tool_input):
    if name == "search_orders":
        r = requests.get(
            "https://api.internal.example.com/orders",
            params={"customer_id": tool_input["customer_id"],
                    "status": tool_input.get("status", "any")},
            headers={"Authorization": f"Bearer {API_TOKEN}"},
            timeout=10,
        )
        r.raise_for_status()
        return r.text  # return only the fields the agent needs

A database tool — never hand the model raw SQL execution against write credentials. Expose parameterized, read-only queries:

import sqlite3

def execute_tool(name, tool_input):
    if name == "count_signups_since":
        con = sqlite3.connect("file:app.db?mode=ro", uri=True)  # read-only
        cur = con.execute(
            "SELECT COUNT(*) FROM users WHERE created_at >= ?",
            (tool_input["since_date"],),
        )
        return str(cur.fetchone()[0])

A file-system tool — confine paths to a sandbox root so the model cannot escape with ../:

import os

SANDBOX = os.path.realpath("./workspace")

def read_file(path):
    full = os.path.realpath(os.path.join(SANDBOX, path))
    if not full.startswith(SANDBOX + os.sep):
        return {"error": "Path escapes the sandbox", "is_error": True}
    with open(full) as f:
        return f.read()
WARNING
Least privilege at the connection layer
Give the agent a read-only database user, a scoped API token, and a sandboxed directory. The schema is a suggestion; the credentials are the real enforcement. Assume the model will eventually request something it shouldn't.
TIP
Return high-signal results
Trim responses to the fields the agent reasons over, and prefer stable identifiers (slugs, UUIDs) over opaque internal references. Bloated results waste context and bury the signal the model needs next.

5 MCP — A Standard Interface for Tools

Hand-writing execute_tool works, but every host re-implements the same plumbing. The Model Context Protocol (MCP) is an open standard that decouples tool providers from consumers: a server exposes capabilities once, and any MCP-aware host (Claude Code, an IDE, your own agent) can use them without bespoke glue.

MCP uses JSON-RPC 2.0 over stateful connections, with two common transports — stdio (a local subprocess) and streamable HTTP (a remote service). A server can expose three primitives:

Primitive Controlled by Example
Tools The model query_database, create_ticket
Resources The app/user A file, a DB schema, an API doc
Prompts The user A templated workflow Claude can run

A host discovers tools with a tools/list request and invokes them with tools/call — the same name/description/JSON-Schema shape from Step 2, delivered over the wire. Registering a server is configuration, not code:

{
  "mcpServers": {
    "orders": {
      "command": "npx",
      "args": ["-y", "@example/mcp-server-orders"],
      "env": { "ORDERS_API_TOKEN": "scoped-read-token" }
    }
  }
}
NOTE
Why MCP matters for agents
Write a tool once as an MCP server and it works across every compatible client. Instead of N hosts times M integrations, you build M servers and N hosts speak one protocol.
WARNING
Human in the loop is in the spec
The MCP specification states there SHOULD always be a human able to deny tool invocations, because a tool description is untrusted by default and a tool is arbitrary code execution. Treat any third-party server like an unvetted dependency.

6 Handling Tool Failures Gracefully

Tools fail constantly — networks time out, parameters are wrong, rate limits hit. The model can recover only if you tell it what went wrong. Return the failure as a tool_result with is_error set to true and an instructive message; never crash the loop.

{
  "role": "user",
  "content": [
    {
      "type": "tool_result",
      "tool_use_id": "toolu_01A09q90qw90lq917835lq9",
      "content": "Rate limit exceeded. Retry after 60 seconds.",
      "is_error": true
    }
  ]
}

Wrap execution so any exception becomes a structured error the model can act on:

def safe_execute(block):
    try:
        output = execute_tool(block.name, block.input)
        return {"type": "tool_result", "tool_use_id": block.id, "content": output}
    except Exception as e:
        return {
            "type": "tool_result",
            "tool_use_id": block.id,
            "content": f"{type(e).__name__}: {e}. Check inputs and retry once.",
            "is_error": True,
        }

A generic "failed" tells the model nothing; "ConnectionError: weather API unavailable (HTTP 500). Try a different city or report the outage." lets it adapt. The API even retries an invalid tool call a few times on its own when the error names the missing piece.

Four design rules that compound across a tool set:

  • Consolidate related operations. Prefer one manage_pr tool with an action parameter over create_pr, review_pr, merge_pr — fewer, richer tools reduce selection ambiguity.
  • Namespace by service when tools span systems: github_list_prs, chat_send_message.
  • Force calls when you must. tool_choice accepts auto (default), any (some tool), tool (a named tool), or none.
  • Eliminate malformed inputs by setting strict: true on definitions, which guarantees inputs match your schema.
TIP
Errors are instructions, not just logs
Every error string you return is a prompt the model reads next. Write them as guidance: name the failure, name the likely cause, and name the recovery action.

Questions & Answers

Q: If a tool description is untrusted and the model picks tools, what stops it from deleting my production database?
The schema does not — credentials do. Connect tools with read-only DB users, scoped tokens, and sandboxed paths, and gate destructive actions behind a human-approval step. The model emits a request; your code decides whether to honor it. Full guardrail patterns are Lesson 7.
Q: How many tools can I attach before the model gets confused or costs spike?
Every tool's name, description, and schema is sent on every request, so a large set inflates input tokens and selection ambiguity. Consolidate related actions into fewer parameterized tools, namespace by service, and attach only what a task needs. For very large libraries, the platform offers a tool-search mechanism that loads definitions on demand.
Q: Should I write a custom execute_tool loop or adopt MCP?
Use a direct loop when the tools live inside one application you fully control — it's the least overhead. Reach for MCP when the same tools must be shared across multiple hosts, exposed to third parties, or run out-of-process for isolation. They compose: an MCP host can still wrap the JSON-RPC calls behind the same tool_use loop you wrote in Step 3.
Q: What happens if the model calls a tool with a parameter I never defined?
With ordinary tools the model can occasionally hallucinate or omit parameters; return a tool_result with is_error true naming the bad field and it will usually retry correctly. To eliminate the failure mode entirely, set strict: true on the definition, which forces inputs to conform to the schema.
Q: Tool calls add a full API round trip each. How do I keep latency sane?
Return multiple tool_use blocks per turn so independent calls run in parallel, keep tool_result payloads lean, and cache stable tool definitions and system prompts so they aren't re-billed every turn. Beyond that, prefer fewer, higher-level tools so the agent reaches its goal in fewer steps.

Key Takeaways

  1. Tools are the contract, not the code. A tool is name + description + input_schema; the model emits requests and your code executes them, keeping every side effect under your control.
  2. The description does the heavy lifting. It is the model's only guide to when and how to call a tool — write several specific sentences, not a fragment.
  3. The loop is ReAct made real. Detect stop_reason: "tool_use", execute, return a tool_result keyed by tool_use_id, and repeat under a bounded iteration cap.
  4. Enforce safety at the boundary. Read-only users, scoped tokens, and sandboxed paths matter more than the schema, because the schema is a suggestion and the credentials are the wall.
  5. MCP standardizes the interface. Build a tool once as an MCP server (JSON-RPC over stdio or HTTP) and any compatible host can use it — with a human able to deny invocations.
  6. Errors are prompts. Return failures as is_error results with instructive text so the agent can recover instead of stalling.

Next Steps: Lesson 4: Memory Systems