Troubleshooting
A field guide to the failure modes you hit running an agent loop in production. Each entry follows the same shape: symptom, cause, fix. Pair it with the tool reference and the rules when debugging.
Quick Triage
| Symptom | Cause | Jump to |
|---|---|---|
| Agent never returns | No stop condition / step cap | Infinite loops |
| Calls a non-existent tool | Vague tool descriptions | Hallucinated calls |
| Repeats the same action | Tool result not fed back | Ignoring results |
max_tokens / context errors |
Transcript grew unbounded | Context overflow |
| "I cannot proceed" | Missing tool / ambiguous goal | Getting stuck |
| Deletes files / overspends | No guardrails | Unsafe actions |
Infinite Loops
The agent calls tools forever, oscillates between two actions, or re-plans without converging.
| Cause | Fix |
|---|---|
| No hard cap on iterations | Enforce a max_steps counter; return partial result when hit |
| No termination signal | Require a finish tool or stop reason to end the loop |
| Same error each turn | Detect repeated calls; inject a corrective message |
| Re-planning never narrows | Track plan diffs; abort if N plans are unchanged |
A correct loop has three exits: a final answer, a step cap, and an error budget.
MAX_STEPS = 12
for step in range(MAX_STEPS):
resp = client.messages.create(model=MODEL, messages=messages, tools=TOOLS)
if resp.stop_reason != "tool_use":
return resp # natural completion -> exit
messages.append({"role": "assistant", "content": resp.content})
messages.append({"role": "user", "content": run_tools(resp.content)})
else:
raise RuntimeError(f"agent exceeded {MAX_STEPS} steps")
Detect oscillation by hashing each (tool_name, input) and bailing on repeats:
key = (block.name, json.dumps(block.input, sort_keys=True))
seen[key] += 1
if seen[key] > 2:
raise RuntimeError(f"looping on {block.name}")
Hallucinated Tool Calls
The model invents a tool name you never defined, or passes arguments that violate the schema.
| Cause | Fix |
|---|---|
Tool name not in the tools list |
Validate block.name against your registry; return an error result, do not crash |
| Description too vague | Write descriptions that state when to use the tool and what it returns |
Loose input_schema |
Add required, enum, and type constraints |
| Asking for capabilities you did not provide | List available tools in the system prompt |
Never let an unknown tool throw. Return a structured error so the model can recover on the next turn:
def dispatch(block):
if block.name not in REGISTRY:
return {"type": "tool_result", "tool_use_id": block.id,
"content": f"Error: unknown tool '{block.name}'.",
"is_error": True}
return REGISTRY[block.name](block.input)
Tight schemas with required, enum, and minimum constraints prevent most argument hallucinations — give the model less room to improvise.
Ignoring Tool Results
The agent calls a tool, then acts as if it never ran — repeating the call, or answering from stale assumptions.
| Cause | Fix |
|---|---|
tool_result not appended to messages |
Always append the assistant turn and the matching tool_result before the next call |
tool_use_id mismatch |
Echo the exact id from each tool_use block in its tool_result |
| Result too large / truncated | Summarise or paginate before returning it |
| Result unparseable | Return clean JSON or plain text, not raw stack traces |
Every tool_use block must be answered by exactly one tool_result with the same id, in the very next user turn:
results = []
for block in resp.content:
if block.type == "tool_use":
results.append({
"type": "tool_result",
"tool_use_id": block.id, # MUST match
"content": str(dispatch(block))[:8000] # cap size
})
messages.append({"role": "user", "content": results})
Context Overflow
The request fails with a token/context error, or the agent "forgets" early instructions deep into a long run.
| Cause | Fix |
|---|---|
| Full transcript resent every turn | Trim or summarise old turns; keep system prompt + recent N turns |
| Verbose tool outputs accumulate | Store large blobs externally, pass back an ID/summary |
| No token accounting | Track usage.input_tokens each turn; compact before the limit |
| Important facts buried mid-history | Pin key state into the system prompt or a scratchpad |
A simple sliding window plus a running summary keeps context bounded:
def compact(messages, keep=6):
if len(messages) <= keep + 1:
return messages
summary = summarise(messages[1:-keep]) # one cheap LLM call
note = {"role": "user", "content": f"[Summary so far]\n{summary}"}
return messages[:1] + [note] + messages[-keep:]
Prefer returning handles over payloads — write the 5 MB file to disk, return the path.
Getting Stuck
The agent declares it cannot continue, asks the same clarifying question repeatedly, or produces empty turns.
| Cause | Fix |
|---|---|
| Required tool is missing | Audit the goal against the tool set; add the missing capability |
| Goal is ambiguous | Force a planning step that restates the goal and lists sub-tasks |
| Silent tool failure | Make tools return explicit errors, not empty strings |
| No recovery path | On repeated failure, switch strategy or escalate to a human |
Give the model an explicit escape hatch — an ask_human tool it can call when blocked — so "stuck" becomes a clean, observable outcome rather than a silent hang. See Lesson 5 for re-planning patterns when the environment shifts mid-task.
Unsafe Actions
The agent deletes data, sends real emails, runs destructive shell commands, or burns through budget unchecked.
| Cause | Fix |
|---|---|
| No human-in-the-loop | Gate write/delete/spend tools behind an approval check |
| Full filesystem / network access | Sandbox tools in a least-privilege container |
| No spend cap | Track cumulative tokens and calls; abort past budget |
| Prompt injection from tool output | Treat tool output as untrusted data, never instructions |
Classify tools by blast radius and require confirmation for the dangerous ones:
DESTRUCTIVE = {"delete_file", "send_email", "run_shell", "charge_card"}
def guarded(block):
if block.name in DESTRUCTIVE and not approve(block):
return {"type": "tool_result", "tool_use_id": block.id,
"content": "Denied by guardrail: action not approved.",
"is_error": True}
return dispatch(block)
Enforce a hard resource budget (cumulative tool calls and tokens) alongside the step cap from Infinite loops — abort the run the moment it is exhausted. Full safety architecture, sandboxing, and the principal hierarchy are covered in Lesson 7.
Where to Go Next
| Need | Resource |
|---|---|
| Tool schemas and loop API | Tool reference |
| Operating conventions | Rules |
| Term definitions | Glossary |
| Multi-agent failures | Agent Orchestration |