Troubleshooting

Reference intermediate

A field guide to the failure modes you hit running an agent loop in production. Each entry follows the same shape: symptom, cause, fix. Pair it with the tool reference and the rules when debugging.

Quick Triage

Symptom Cause Jump to
Agent never returns No stop condition / step cap Infinite loops
Calls a non-existent tool Vague tool descriptions Hallucinated calls
Repeats the same action Tool result not fed back Ignoring results
max_tokens / context errors Transcript grew unbounded Context overflow
"I cannot proceed" Missing tool / ambiguous goal Getting stuck
Deletes files / overspends No guardrails Unsafe actions

Infinite Loops

The agent calls tools forever, oscillates between two actions, or re-plans without converging.

Cause Fix
No hard cap on iterations Enforce a max_steps counter; return partial result when hit
No termination signal Require a finish tool or stop reason to end the loop
Same error each turn Detect repeated calls; inject a corrective message
Re-planning never narrows Track plan diffs; abort if N plans are unchanged

A correct loop has three exits: a final answer, a step cap, and an error budget.

MAX_STEPS = 12
for step in range(MAX_STEPS):
    resp = client.messages.create(model=MODEL, messages=messages, tools=TOOLS)
    if resp.stop_reason != "tool_use":
        return resp  # natural completion -> exit
    messages.append({"role": "assistant", "content": resp.content})
    messages.append({"role": "user", "content": run_tools(resp.content)})
else:
    raise RuntimeError(f"agent exceeded {MAX_STEPS} steps")

Detect oscillation by hashing each (tool_name, input) and bailing on repeats:

key = (block.name, json.dumps(block.input, sort_keys=True))
seen[key] += 1
if seen[key] > 2:
    raise RuntimeError(f"looping on {block.name}")

Hallucinated Tool Calls

The model invents a tool name you never defined, or passes arguments that violate the schema.

Cause Fix
Tool name not in the tools list Validate block.name against your registry; return an error result, do not crash
Description too vague Write descriptions that state when to use the tool and what it returns
Loose input_schema Add required, enum, and type constraints
Asking for capabilities you did not provide List available tools in the system prompt

Never let an unknown tool throw. Return a structured error so the model can recover on the next turn:

def dispatch(block):
    if block.name not in REGISTRY:
        return {"type": "tool_result", "tool_use_id": block.id,
                "content": f"Error: unknown tool '{block.name}'.",
                "is_error": True}
    return REGISTRY[block.name](block.input)

Tight schemas with required, enum, and minimum constraints prevent most argument hallucinations — give the model less room to improvise.

Ignoring Tool Results

The agent calls a tool, then acts as if it never ran — repeating the call, or answering from stale assumptions.

Cause Fix
tool_result not appended to messages Always append the assistant turn and the matching tool_result before the next call
tool_use_id mismatch Echo the exact id from each tool_use block in its tool_result
Result too large / truncated Summarise or paginate before returning it
Result unparseable Return clean JSON or plain text, not raw stack traces

Every tool_use block must be answered by exactly one tool_result with the same id, in the very next user turn:

results = []
for block in resp.content:
    if block.type == "tool_use":
        results.append({
            "type": "tool_result",
            "tool_use_id": block.id,        # MUST match
            "content": str(dispatch(block))[:8000]  # cap size
        })
messages.append({"role": "user", "content": results})

Context Overflow

The request fails with a token/context error, or the agent "forgets" early instructions deep into a long run.

Cause Fix
Full transcript resent every turn Trim or summarise old turns; keep system prompt + recent N turns
Verbose tool outputs accumulate Store large blobs externally, pass back an ID/summary
No token accounting Track usage.input_tokens each turn; compact before the limit
Important facts buried mid-history Pin key state into the system prompt or a scratchpad

A simple sliding window plus a running summary keeps context bounded:

def compact(messages, keep=6):
    if len(messages) <= keep + 1:
        return messages
    summary = summarise(messages[1:-keep])  # one cheap LLM call
    note = {"role": "user", "content": f"[Summary so far]\n{summary}"}
    return messages[:1] + [note] + messages[-keep:]

Prefer returning handles over payloads — write the 5 MB file to disk, return the path.

Getting Stuck

The agent declares it cannot continue, asks the same clarifying question repeatedly, or produces empty turns.

Cause Fix
Required tool is missing Audit the goal against the tool set; add the missing capability
Goal is ambiguous Force a planning step that restates the goal and lists sub-tasks
Silent tool failure Make tools return explicit errors, not empty strings
No recovery path On repeated failure, switch strategy or escalate to a human

Give the model an explicit escape hatch — an ask_human tool it can call when blocked — so "stuck" becomes a clean, observable outcome rather than a silent hang. See Lesson 5 for re-planning patterns when the environment shifts mid-task.

Unsafe Actions

The agent deletes data, sends real emails, runs destructive shell commands, or burns through budget unchecked.

Cause Fix
No human-in-the-loop Gate write/delete/spend tools behind an approval check
Full filesystem / network access Sandbox tools in a least-privilege container
No spend cap Track cumulative tokens and calls; abort past budget
Prompt injection from tool output Treat tool output as untrusted data, never instructions

Classify tools by blast radius and require confirmation for the dangerous ones:

DESTRUCTIVE = {"delete_file", "send_email", "run_shell", "charge_card"}

def guarded(block):
    if block.name in DESTRUCTIVE and not approve(block):
        return {"type": "tool_result", "tool_use_id": block.id,
                "content": "Denied by guardrail: action not approved.",
                "is_error": True}
    return dispatch(block)

Enforce a hard resource budget (cumulative tool calls and tokens) alongside the step cap from Infinite loops — abort the run the moment it is exhausted. Full safety architecture, sandboxing, and the principal hierarchy are covered in Lesson 7.

Where to Go Next

Need Resource
Tool schemas and loop API Tool reference
Operating conventions Rules
Term definitions Glossary
Multi-agent failures Agent Orchestration