Tool Use — Giving Agents Hands
Learning Outcomes
- Define tools with correct JSON schemas that models can select and parameterize reliably
- Implement the tool-use loop: tool_use stop reason, execution, and tool_result return
- Connect agents to REST APIs, databases, and file systems behind well-designed tools
- Use the Model Context Protocol (MCP) to expose tools through a standard interface
- Handle tool failures gracefully with instructive error messages and validation
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why tools turn LLMs into agents |
| Concept | 6 min | The function-calling contract |
| Build | 8 min | Defining a tool schema |
| Build | 8 min | The agent execution loop |
| Build | 7 min | Wrapping APIs, databases, and files |
| Concept | 6 min | MCP — the standard interface |
| Hardening | 5 min | Error handling and tool design |
| Wrap-up | 2 min | Key takeaways, next lesson |
Before You Begin
Pre-work:
- Complete Lesson 1: What Are AI Agents? and Lesson 2: Agent Architectures — this lesson assumes you understand the ReAct loop.
- Have an Anthropic API key exported as
ANTHROPIC_API_KEY. - Be comfortable reading JSON Schema and Python.
Shopping List:
- Python 3.10+ and the official SDK:
pip install anthropic - A terminal with
node/npxavailable (for the MCP demo) - Optional: a SQLite or PostgreSQL database to wrap as a tool
A bare LLM can only emit text. Tool use (also called function calling) lets a model ask your code to run a function, then read the result back. The model never executes anything itself — it produces a structured request, your code runs it, and you feed the output back into the conversation.
The Anthropic Messages API splits tools into client tools that run in your application, and server tools (like web_search and code_execution) that run on Anthropic's infrastructure and return results directly. This lesson focuses on client tools — that is where you give an agent hands on your systems.
The full round trip has four moves:
1. You send: messages + a `tools` list (each tool = name + description + input_schema)
2. Claude replies with stop_reason "tool_use" and a tool_use block (id, name, input)
3. You execute the function and send back a tool_result block (matched by tool_use_id)
4. Claude reads the result and either calls another tool or answers
Steps 2 through 4 repeat until the model stops requesting tools — that loop is the ReAct architecture from Lesson 2, made concrete.
A client tool is defined by three fields: a name (matching ^[a-zA-Z0-9_-]{1,64}$), a plaintext description, and an input_schema that is a standard JSON Schema object. The canonical shape:
{
"name": "get_weather",
"description": "Get the current weather in a given location. Use this when the user asks about current conditions or temperature for a specific city. Returns temperature and a short summary. Does not return forecasts.",
"input_schema": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, CA"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit; defaults to celsius"
}
},
"required": ["location"]
}
}
The description is the highest-leverage thing you write: it is the model's only signal for deciding whether to call the tool and how to fill the parameters. Aim for several sentences covering what it does, when to use it (and when not to), and what each parameter means.
| Field | Required | Purpose |
|---|---|---|
name |
Yes | Stable identifier the model emits to call the tool |
description |
Yes | The model's only guide to selection and arguments |
input_schema |
Yes | JSON Schema; constrains and documents inputs |
strict |
No | Set true to guarantee inputs match the schema exactly |
description to every property and an enum wherever values are constrained. The model treats the schema as a spec — a vague schema produces vague tool calls.Gets the price for a ticker leaves the model guessing about format, scope, and timing. Under-specified tools are the number-one cause of wrong or skipped tool calls.Now wire the contract into a loop. When Claude wants a tool, the response's stop_reason is "tool_use" and it contains one or more tool_use blocks, each with an id, name, and input. You run the matching function and reply with a user message whose content is a tool_result block keyed by tool_use_id.
import anthropic
client = anthropic.Anthropic()
TOOLS = [ { "name": "get_weather", "description": "...", "input_schema": {...} } ]
def execute_tool(name, tool_input):
if name == "get_weather":
# call your real weather code here
return f"18 degrees, partly cloudy in {tool_input['location']}"
raise ValueError(f"Unknown tool: {name}")
messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]
while True:
resp = client.messages.create(
model="claude-opus-4-8",
max_tokens=1024,
tools=TOOLS,
messages=messages,
)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use":
break # Claude produced a final answer
results = []
for block in resp.content:
if block.type == "tool_use":
output = execute_tool(block.name, block.input)
results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": output,
})
messages.append({"role": "user", "content": results})
print(resp.content[-1].text)
Two formatting rules the API enforces: tool_result blocks must come first in the user message's content array, and they must immediately follow the assistant turn that requested them — no messages in between.
tool_use blocks. The loop handles this by iterating over every block and returning one tool_result per call before the next round trip.Tools are just functions, so anything callable from your runtime can become one. The art is in the boundary you expose.
A REST API tool — wrap the call, not the raw endpoint:
import requests
def execute_tool(name, tool_input):
if name == "search_orders":
r = requests.get(
"https://api.internal.example.com/orders",
params={"customer_id": tool_input["customer_id"],
"status": tool_input.get("status", "any")},
headers={"Authorization": f"Bearer {API_TOKEN}"},
timeout=10,
)
r.raise_for_status()
return r.text # return only the fields the agent needs
A database tool — never hand the model raw SQL execution against write credentials. Expose parameterized, read-only queries:
import sqlite3
def execute_tool(name, tool_input):
if name == "count_signups_since":
con = sqlite3.connect("file:app.db?mode=ro", uri=True) # read-only
cur = con.execute(
"SELECT COUNT(*) FROM users WHERE created_at >= ?",
(tool_input["since_date"],),
)
return str(cur.fetchone()[0])
A file-system tool — confine paths to a sandbox root so the model cannot escape with ../:
import os
SANDBOX = os.path.realpath("./workspace")
def read_file(path):
full = os.path.realpath(os.path.join(SANDBOX, path))
if not full.startswith(SANDBOX + os.sep):
return {"error": "Path escapes the sandbox", "is_error": True}
with open(full) as f:
return f.read()
Hand-writing execute_tool works, but every host re-implements the same plumbing. The Model Context Protocol (MCP) is an open standard that decouples tool providers from consumers: a server exposes capabilities once, and any MCP-aware host (Claude Code, an IDE, your own agent) can use them without bespoke glue.
MCP uses JSON-RPC 2.0 over stateful connections, with two common transports — stdio (a local subprocess) and streamable HTTP (a remote service). A server can expose three primitives:
| Primitive | Controlled by | Example |
|---|---|---|
| Tools | The model | query_database, create_ticket |
| Resources | The app/user | A file, a DB schema, an API doc |
| Prompts | The user | A templated workflow Claude can run |
A host discovers tools with a tools/list request and invokes them with tools/call — the same name/description/JSON-Schema shape from Step 2, delivered over the wire. Registering a server is configuration, not code:
{
"mcpServers": {
"orders": {
"command": "npx",
"args": ["-y", "@example/mcp-server-orders"],
"env": { "ORDERS_API_TOKEN": "scoped-read-token" }
}
}
}
Tools fail constantly — networks time out, parameters are wrong, rate limits hit. The model can recover only if you tell it what went wrong. Return the failure as a tool_result with is_error set to true and an instructive message; never crash the loop.
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_01A09q90qw90lq917835lq9",
"content": "Rate limit exceeded. Retry after 60 seconds.",
"is_error": true
}
]
}
Wrap execution so any exception becomes a structured error the model can act on:
def safe_execute(block):
try:
output = execute_tool(block.name, block.input)
return {"type": "tool_result", "tool_use_id": block.id, "content": output}
except Exception as e:
return {
"type": "tool_result",
"tool_use_id": block.id,
"content": f"{type(e).__name__}: {e}. Check inputs and retry once.",
"is_error": True,
}
A generic "failed" tells the model nothing; "ConnectionError: weather API unavailable (HTTP 500). Try a different city or report the outage." lets it adapt. The API even retries an invalid tool call a few times on its own when the error names the missing piece.
Four design rules that compound across a tool set:
- Consolidate related operations. Prefer one
manage_prtool with anactionparameter overcreate_pr,review_pr,merge_pr— fewer, richer tools reduce selection ambiguity. - Namespace by service when tools span systems:
github_list_prs,chat_send_message. - Force calls when you must.
tool_choiceacceptsauto(default),any(some tool),tool(a named tool), ornone. - Eliminate malformed inputs by setting
strict: trueon definitions, which guarantees inputs match your schema.
Questions & Answers
strict: true on the definition, which forces inputs to conform to the schema.Key Takeaways
- Tools are the contract, not the code. A tool is
name+description+input_schema; the model emits requests and your code executes them, keeping every side effect under your control. - The description does the heavy lifting. It is the model's only guide to when and how to call a tool — write several specific sentences, not a fragment.
- The loop is ReAct made real. Detect
stop_reason: "tool_use", execute, return atool_resultkeyed bytool_use_id, and repeat under a bounded iteration cap. - Enforce safety at the boundary. Read-only users, scoped tokens, and sandboxed paths matter more than the schema, because the schema is a suggestion and the credentials are the wall.
- MCP standardizes the interface. Build a tool once as an MCP server (JSON-RPC over stdio or HTTP) and any compatible host can use it — with a human able to deny invocations.
- Errors are prompts. Return failures as
is_errorresults with instructive text so the agent can recover instead of stalling.
Next Steps: Lesson 4: Memory Systems