Planning & Reasoning
Learning Outcomes
- Apply chain-of-thought prompting to make a model's reasoning explicit and inspectable
- Implement goal decomposition that breaks an objective into an executable task graph
- Compare tree-of-thought search against linear reasoning and choose the right one per task
- Validate a generated plan before execution so bad plans fail fast and cheap
- Build a re-planning loop that reacts to tool failures and environment changes
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why planning separates agents from tool-callers |
| Explain | 6 min | Chain-of-thought and structured reasoning |
| Demo | 8 min | Goal decomposition into a task graph |
| Demo | 7 min | Tree-of-thought and self-evaluation |
| Demo | 7 min | Validating a plan before execution |
| Demo | 8 min | The re-plan loop |
| Wrap-up | 6 min | Failure modes, takeaways |
Before You Begin
Pre-work:
- Complete Lesson 2: Agent Architectures — recognise ReAct and Plan-and-Execute
- Complete Lesson 3: Tool Use — planning produces actions, and actions are tool calls
- Skim Lesson 4: Memory Systems — plans live in working memory
Shopping List:
- Python 3.10+ with the
anthropicSDK (pip install anthropic) - An
ANTHROPIC_API_KEYset - A scratch directory; no external services needed
Chain-of-thought (CoT) prompting asks the model to write out its reasoning before answering — scratch space that doubles as an audit trail. For agents, keep reasoning separate from the answer so your code can route each:
import anthropic
client = anthropic.Anthropic()
SYSTEM = """Reason through the problem inside <thinking> tags,
then give your final answer inside <answer> tags."""
resp = client.messages.create(model="claude-sonnet-4-6", max_tokens=1024,
system=SYSTEM, messages=[{"role": "user", "content":
"3 invoices are overdue, 1 is paid. Which need a reminder, in what order?"}])
Your loop parses the tags, logs <thinking>, and feeds only <answer> onward. Recent Claude models also expose native extended thinking — a separate thinking block.
A real goal — "research the top three vector databases and write a comparison" — is not one action. Goal decomposition turns it into sub-tasks with dependencies. Your planner prompt should demand JSON: a goal and a tasks list, each task with an id, a one-action description, a tool, and depends_on.
The result is a directed acyclic graph (DAG). The executor keeps a done set and each pass runs the ready tasks — those whose dependencies are done:
def ready_tasks(tasks, done):
return [t for t in tasks if t["id"] not in done
and all(dep in done for dep in t["depends_on"])]
If ready_tasks returns empty before all tasks finish, the graph has a cycle or missing edge — fail fast.
Chain-of-thought is a single line of reasoning. When the first approach is often wrong — puzzles, ambiguous specs, anything needing backtracking — tree-of-thought (ToT) explores several branches, scores them, and expands the best. Two helpers do the work — candidates(state, n) returns next steps, score(state, step) rates each 0-10 — feeding a beam search:
def tree_of_thought(root, depth=3, beam_width=2):
frontier = [root]
for _ in range(depth):
scored = [(score(s, c), s + "\n- " + c)
for s in frontier for c in candidates(s)]
scored.sort(reverse=True, key=lambda x: x[0])
frontier = [s for _, s in scored[:beam_width]] # keep top branches
return frontier[0]
ToT is far more expensive — cost grows with depth * beam_width * candidates. Weigh the alternatives:
| Technique | Calls per task | Backtracking | Best for |
|---|---|---|---|
| Chain-of-thought | 1 | No | Linear, well-defined tasks |
| Self-consistency (sample N) | N | No | Verifiable answers |
| Tree-of-thought | Many | Yes | Search, puzzles, ambiguity |
The cheapest failure is one you catch before any tool runs. Gate the plan through structural validation (deterministic code) then semantic validation (a critic) — the free structural check first:
ALLOWED = {"web_search", "read_file", "summarize", "write_file"}
def validate_structure(graph):
tasks = graph.get("tasks") or []
errors, seen = [], set()
for t in tasks:
if t["tool"] not in ALLOWED:
errors.append(f"{t['id']}: unknown tool {t['tool']!r}")
for dep in t["depends_on"]:
if dep not in seen: # missing, or defined later (cycle risk)
errors.append(f"{t['id']}: bad dependency {dep!r}")
seen.add(t["id"])
return errors or (["Plan has no tasks"] if not tasks else [])
If errors come back, feed them to the planner for repair before calling the critic. Only structurally valid plans deserve semantic review — a second model asked does this plan achieve the goal, with no redundant or destructive steps?, returning an approved boolean.
Plan-and-Execute commits upfront. ReAct instead interleaves one reasoning step with one action, observes the result, then decides the next move — re-planning implicitly every turn. The model emits a tool_use block, you run it and append a tool_result — reason → act → observe, capped by steps:
def run_react(goal, tools, max_steps=8):
messages = [{"role": "user", "content": goal}]
for _ in range(max_steps):
resp = client.messages.create(model="claude-sonnet-4-6", max_tokens=1024,
tools=tools, messages=messages)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use":
return resp.content # model decided it is done
results = [{"type": "tool_result", "tool_use_id": b.id,
"content": execute_tool(b.name, b.input)}
for b in resp.content if b.type == "tool_use"]
messages.append({"role": "user", "content": results})
raise RuntimeError("ReAct exceeded max_steps")
The trade-off versus a pre-built plan is latency: ReAct serialises model calls, where a plan runs independent tasks in parallel. Lesson 6 builds an agent on it.
The environment will not cooperate: a search returns nothing, a file is missing, an API rate-limits you. A plan made before observing the world will be partially wrong, so re-planning — revising the plan when it no longer fits reality — defines a real agent. On failure, hand the planner the goal, what succeeded, and the error, and ask for a revised remainder:
def execute_with_replanning(goal, plan, max_replans=3):
done, results = set(), {}
for attempt in range(max_replans + 1):
try:
return run_plan(plan, done, results) # the Step 2 executor
except ToolFailure as err:
if attempt == max_replans:
raise
ctx = {"goal": goal, "completed": list(done), "failure": str(err),
"results": {k: str(v)[:200] for k, v in results.items()}}
reply = client.messages.create(model="claude-sonnet-4-6", max_tokens=1500,
system=REPLAN_SYSTEM, messages=[{"role": "user", "content": json.dumps(ctx)}])
plan = json.loads(reply.content[0].text) # revised remainder
raise RuntimeError("Exhausted re-planning budget")
The context — goal, completed ids with truncated results, and the failure — lets the new plan avoid drifting, redoing work, or repeating itself: a different tool, a broader query, a repair step, or escalation — not a retry.
Questions & Answers
input_schema so the model is constrained to your shape, validate structurally and feed errors back for repair (Step 4), and bound the loop. That eliminates most drift.Key Takeaways
- Make reasoning explicit and structured — separate thinking from the answer so you can log, validate, and debug what the agent does, not just its conclusion.
- Decompose goals into a task graph — a dependency DAG of single-tool tasks; the ready frontier drives execution and parallelism.
- Reach for tree-of-thought only when it pays — chain-of-thought handles most work; reserve ToT for where early commitment is costly.
- Validate plans before executing them — cheap structural checks first, then a critic, with a bounded repair loop. The cheapest failure is caught before a tool runs.
- Bound every loop and re-plan when reality disagrees — step caps, re-plan budgets, and cost ceilings are mandatory; on failure, revise the strategy rather than blindly retrying.
Next Steps: Lesson 6: Building Your First Agent