Planning & Reasoning

45 min intermediate Lesson 5

Learning Outcomes

  • Apply chain-of-thought prompting to make a model's reasoning explicit and inspectable
  • Implement goal decomposition that breaks an objective into an executable task graph
  • Compare tree-of-thought search against linear reasoning and choose the right one per task
  • Validate a generated plan before execution so bad plans fail fast and cheap
  • Build a re-planning loop that reacts to tool failures and environment changes

Lesson Plan

Segment Duration Topic
Intro 3 min Why planning separates agents from tool-callers
Explain 6 min Chain-of-thought and structured reasoning
Demo 8 min Goal decomposition into a task graph
Demo 7 min Tree-of-thought and self-evaluation
Demo 7 min Validating a plan before execution
Demo 8 min The re-plan loop
Wrap-up 6 min Failure modes, takeaways

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ with the anthropic SDK (pip install anthropic)
  • An ANTHROPIC_API_KEY set
  • A scratch directory; no external services needed

1 Reasoning Before Acting: Chain-of-Thought

Chain-of-thought (CoT) prompting asks the model to write out its reasoning before answering — scratch space that doubles as an audit trail. For agents, keep reasoning separate from the answer so your code can route each:

import anthropic
client = anthropic.Anthropic()

SYSTEM = """Reason through the problem inside <thinking> tags,
then give your final answer inside <answer> tags."""

resp = client.messages.create(model="claude-sonnet-4-6", max_tokens=1024,
    system=SYSTEM, messages=[{"role": "user", "content":
        "3 invoices are overdue, 1 is paid. Which need a reminder, in what order?"}])

Your loop parses the tags, logs <thinking>, and feeds only <answer> onward. Recent Claude models also expose native extended thinking — a separate thinking block.

NOTE
Why structure beats 'think step by step'
Tagged sections let you log reasoning separately, validate the answer alone, and strip thinking from memory to save tokens — reasoning as a first-class output.

2 Goal Decomposition: From Objective to Task Graph

A real goal — "research the top three vector databases and write a comparison" — is not one action. Goal decomposition turns it into sub-tasks with dependencies. Your planner prompt should demand JSON: a goal and a tasks list, each task with an id, a one-action description, a tool, and depends_on.

The result is a directed acyclic graph (DAG). The executor keeps a done set and each pass runs the ready tasks — those whose dependencies are done:

def ready_tasks(tasks, done):
    return [t for t in tasks if t["id"] not in done
            and all(dep in done for dep in t["depends_on"])]

If ready_tasks returns empty before all tasks finish, the graph has a cycle or missing edge — fail fast.

WARNING
Plans are hypotheses, not gospel
A decomposed plan is the model's best guess about a world it has not observed — sub-task three may prove unnecessary once sub-task two returns data. Step 6 covers re-planning, the difference between a script and an agent.

3 Tree-of-Thought: Exploring Multiple Paths

Chain-of-thought is a single line of reasoning. When the first approach is often wrong — puzzles, ambiguous specs, anything needing backtracking — tree-of-thought (ToT) explores several branches, scores them, and expands the best. Two helpers do the work — candidates(state, n) returns next steps, score(state, step) rates each 0-10 — feeding a beam search:

def tree_of_thought(root, depth=3, beam_width=2):
    frontier = [root]
    for _ in range(depth):
        scored = [(score(s, c), s + "\n- " + c)
                  for s in frontier for c in candidates(s)]
        scored.sort(reverse=True, key=lambda x: x[0])
        frontier = [s for _, s in scored[:beam_width]]   # keep top branches
    return frontier[0]

ToT is far more expensive — cost grows with depth * beam_width * candidates. Weigh the alternatives:

Technique Calls per task Backtracking Best for
Chain-of-thought 1 No Linear, well-defined tasks
Self-consistency (sample N) N No Verifiable answers
Tree-of-thought Many Yes Search, puzzles, ambiguity
TIP
Start linear, escalate to a tree
Adopt ToT only after you see the agent committing early to bad paths it cannot recover from — the extra calls cost money and latency, so measure the win first.

4 Validating a Plan Before You Execute It

The cheapest failure is one you catch before any tool runs. Gate the plan through structural validation (deterministic code) then semantic validation (a critic) — the free structural check first:

ALLOWED = {"web_search", "read_file", "summarize", "write_file"}

def validate_structure(graph):
    tasks = graph.get("tasks") or []
    errors, seen = [], set()
    for t in tasks:
        if t["tool"] not in ALLOWED:
            errors.append(f"{t['id']}: unknown tool {t['tool']!r}")
        for dep in t["depends_on"]:
            if dep not in seen:          # missing, or defined later (cycle risk)
                errors.append(f"{t['id']}: bad dependency {dep!r}")
        seen.add(t["id"])
    return errors or (["Plan has no tasks"] if not tasks else [])

If errors come back, feed them to the planner for repair before calling the critic. Only structurally valid plans deserve semantic review — a second model asked does this plan achieve the goal, with no redundant or destructive steps?, returning an approved boolean.

WARNING
Validation is your first safety layer — cap the repair loop
Refusing forward edges, unknown tools, and destructive steps before execution prevents runaway behaviour. But planner-critic repair loops forever if the goal is impossible — bound it, then escalate to a human.

5 Interleaving Plan and Act: The ReAct Loop

Plan-and-Execute commits upfront. ReAct instead interleaves one reasoning step with one action, observes the result, then decides the next move — re-planning implicitly every turn. The model emits a tool_use block, you run it and append a tool_result — reason → act → observe, capped by steps:

def run_react(goal, tools, max_steps=8):
    messages = [{"role": "user", "content": goal}]
    for _ in range(max_steps):
        resp = client.messages.create(model="claude-sonnet-4-6", max_tokens=1024,
                                       tools=tools, messages=messages)
        messages.append({"role": "assistant", "content": resp.content})
        if resp.stop_reason != "tool_use":
            return resp.content                      # model decided it is done
        results = [{"type": "tool_result", "tool_use_id": b.id,
                    "content": execute_tool(b.name, b.input)}
                   for b in resp.content if b.type == "tool_use"]
        messages.append({"role": "user", "content": results})
    raise RuntimeError("ReAct exceeded max_steps")

The trade-off versus a pre-built plan is latency: ReAct serialises model calls, where a plan runs independent tasks in parallel. Lesson 6 builds an agent on it.

TIP
max_steps is non-negotiable
Every agent loop needs a hard step cap, or a confused agent calls tools forever, burning tokens and money. The cap is your circuit breaker — set it, log hits, and treat them as bugs.

6 Re-planning When Reality Disagrees

The environment will not cooperate: a search returns nothing, a file is missing, an API rate-limits you. A plan made before observing the world will be partially wrong, so re-planning — revising the plan when it no longer fits reality — defines a real agent. On failure, hand the planner the goal, what succeeded, and the error, and ask for a revised remainder:

def execute_with_replanning(goal, plan, max_replans=3):
    done, results = set(), {}
    for attempt in range(max_replans + 1):
        try:
            return run_plan(plan, done, results)          # the Step 2 executor
        except ToolFailure as err:
            if attempt == max_replans:
                raise
            ctx = {"goal": goal, "completed": list(done), "failure": str(err),
                   "results": {k: str(v)[:200] for k, v in results.items()}}
            reply = client.messages.create(model="claude-sonnet-4-6", max_tokens=1500,
                system=REPLAN_SYSTEM, messages=[{"role": "user", "content": json.dumps(ctx)}])
            plan = json.loads(reply.content[0].text)      # revised remainder
    raise RuntimeError("Exhausted re-planning budget")

The context — goal, completed ids with truncated results, and the failure — lets the new plan avoid drifting, redoing work, or repeating itself: a different tool, a broader query, a repair step, or escalation — not a retry.

WARNING
Re-planning can mask a doomed goal
If your agent re-plans three times and still fails, the problem is usually the goal or the tool set, not the plan. Bound re-plans and fail loudly with the history attached.

Questions & Answers

Q: If the model reasons internally already, why force chain-of-thought?
It improves accuracy because the model attends to its own intermediate tokens. More importantly for agents, it is inspectable — you cannot debug or validate reasoning you cannot see.
Q: Tree-of-thought sounds expensive. When is it worth it?
Only when early commitment is costly and the task supports backtracking — constraint puzzles, many viable orderings, ambiguous specs. For the typical "fetch, transform, summarise" task, chain-of-thought wins.
Q: Plan-and-Execute or ReAct?
Upfront planning gives a validatable, parallelisable DAG and lower latency — good when the task is predictable. ReAct adapts turn-by-turn — good when each observation changes the next step. Many agents are hybrids: a coarse plan, then ReAct per phase.
Q: My planner keeps emitting JSON with subtle schema errors. How do I make it reliable?
Define the plan as a tool with a strict input_schema so the model is constrained to your shape, validate structurally and feed errors back for repair (Step 4), and bound the loop. That eliminates most drift.
Q: How do I stop a re-planning loop from running up a huge bill?
Combine three caps: a re-plan budget, a total step cap, and a hard token/cost ceiling per run. When any trips, halt and surface the history to a human — an alert, not a silent retry. Lesson 7 formalises these.

Key Takeaways

  1. Make reasoning explicit and structured — separate thinking from the answer so you can log, validate, and debug what the agent does, not just its conclusion.
  2. Decompose goals into a task graph — a dependency DAG of single-tool tasks; the ready frontier drives execution and parallelism.
  3. Reach for tree-of-thought only when it pays — chain-of-thought handles most work; reserve ToT for where early commitment is costly.
  4. Validate plans before executing them — cheap structural checks first, then a critic, with a bounded repair loop. The cheapest failure is caught before a tool runs.
  5. Bound every loop and re-plan when reality disagrees — step caps, re-plan budgets, and cost ceilings are mandatory; on failure, revise the strategy rather than blindly retrying.

Next Steps: Lesson 6: Building Your First Agent