Agent Architectures
Learning Outcomes
- Diagram the think-act-observe loop that underpins every agent architecture
- Implement a minimal ReAct loop against Claude's tool use API
- Contrast Plan-and-Execute with ReAct and explain when each wins
- Add a Reflexion critique step to improve an agent's output quality
- Choose an architecture based on latency, cost, and reliability trade-offs
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why architecture matters more than the model |
| Explain | 6 min | The universal think-act-observe loop |
| Build | 8 min | Implementing ReAct with Claude tool use |
| Explain | 7 min | Plan-and-Execute: separating planning from doing |
| Build | 6 min | Adding a Reflexion critique loop |
| Compare | 6 min | Trade-offs and a decision table |
| Wrap-up | 4 min | Key takeaways and next steps |
Before You Begin
Pre-work:
- Complete Lesson 1: What Are AI Agents? so the planning / tool-use / memory vocabulary is fresh
- Skim the glossary for the terms agent loop, tool call, and observation
- Be comfortable making a basic Claude API request in Python
Shopping List:
- Python 3.10+ with the
anthropicSDK installed (pip install anthropic) - An Anthropic API key exported as an environment variable
- A code editor and a terminal
- A scratch directory for the snippets in this lesson
Before we name any architecture, internalise the one structure they all share. An agent is a model placed inside a loop that lets it act, see the result, and decide again. Strip away the branding and every pattern is a variation on think → act → observe → repeat.
THINK ──▶ ACT ──▶ OBSERVE ──┐
(reason) (tool) (result) │
▲ │
└─── repeat until done ◀──┘
│
▼ FINISH (goal met)
The three things that distinguish architectures are: where the thinking happens (all upfront, or interleaved with acting), whether the agent critiques itself, and how it decides to stop. That last point is the one beginners underestimate — a loop with no stop condition is a bug that bills you per token.
Three terms we will reuse throughout:
| Term | Meaning |
|---|---|
| Tool call | The model's structured request to run an external function |
| Observation | The result returned to the model after a tool runs |
| Terminal state | The condition that ends the loop (goal met, budget spent, or step cap hit) |
max_steps as non-negotiable.ReAct (Reasoning + Acting) is the simplest and most common architecture. The model alternates between a reasoning thought and a single action, then sees the observation before deciding the next move. There is no separate planning phase — the plan emerges one step at a time.
Here is the control flow in pseudocode:
function react(goal, tools, max_steps):
messages = [user(goal)]
for step in range(max_steps):
response = model.generate(messages, tools)
if response.is_final_answer:
return response.text
for call in response.tool_calls:
result = run_tool(call.name, call.input)
messages.append(assistant(call))
messages.append(tool_result(call.id, result))
return "Stopped: step limit reached"
The strength is simplicity: the model adapts to whatever the last observation told it. The weakness is that it can wander — without a plan, a ReAct agent on a long task may loop, repeat failed actions, or lose the thread. It is the natural fit for Claude because tool use is the act step: the model requests a tool, you run it, you feed the result back, and the loop continues until the model stops requesting tools.
Let's make it concrete. First, define a tool. Claude's tool use API takes a JSON schema describing each tool's name, purpose, and inputs:
import anthropic
client = anthropic.Anthropic()
TOOLS = [
{
"name": "get_word_length",
"description": "Return the number of characters in a word.",
"input_schema": {
"type": "object",
"properties": {
"word": {"type": "string", "description": "The word to measure"}
},
"required": ["word"],
},
}
]
def run_tool(name, tool_input):
if name == "get_word_length":
return str(len(tool_input["word"]))
return f"Unknown tool: {name}"
Now the loop. Notice the three classic moves: send messages, run any requested tools, append the results, repeat. The terminal state is stop_reason == "end_turn" — Claude has stopped asking for tools.
def react_agent(goal, max_steps=8):
messages = [{"role": "user", "content": goal}]
for step in range(max_steps):
resp = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=1024,
tools=TOOLS,
messages=messages,
)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use":
return "".join(b.text for b in resp.content if b.type == "text")
tool_results = []
for block in resp.content:
if block.type == "tool_use":
output = run_tool(block.name, block.input)
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": output,
})
messages.append({"role": "user", "content": tool_results})
return "Stopped: reached max_steps"
print(react_agent("How many letters are in the word 'architecture'?"))
The model reasons ("I should measure the word"), emits a tool_use block, you run it, append a tool_result, and Claude produces the final text answer on the next turn. That is a complete ReAct agent in about thirty lines.
user role message. The assistant requested the tool; the user (your harness) supplies the observation. Mismatching the roles is the most common first-time bug.ReAct decides one step at a time. Plan-and-Execute splits the work into two phases: a planner produces a full ordered list of steps upfront, then an executor carries out each step (often a small ReAct loop per step). A separate planning pass means the agent commits to a strategy before spending tokens on actions.
GOAL ──▶ PLANNER ──▶ [step1, step2, step3, ...]
│
▼
EXECUTOR ──▶ run step, observe ──▶ RE-PLAN?
▲ │
└──── next step ◀─────────────────┘
You can ask the planner to emit a machine-readable plan so the executor can iterate over it deterministically:
PLAN_PROMPT = """You are a planner. Break the goal into an ordered list of
concrete steps. Respond with ONLY a JSON array of strings, no prose.
Goal: {goal}"""
def make_plan(goal):
resp = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=512,
messages=[{"role": "user",
"content": PLAN_PROMPT.format(goal=goal)}],
)
import json
return json.loads(resp.content[0].text)
# plan = make_plan("Research the top 3 Python web frameworks and compare them")
# for step in plan:
# execute_step(step) # each step can be its own ReAct loop
The big advantage is on complex, multi-step tasks: a plan keeps a long task on track and makes the agent's intent auditable before any side effects occur. The cost is latency and rigidity — you pay for a planning round-trip, and a plan made before seeing reality may be wrong. That is why production Plan-and-Execute systems add a re-plan step: after each executed step, check whether the remaining plan still makes sense, and regenerate it if the world changed.
Reflexion adds a self-improvement loop on top of either pattern. The agent produces a draft, a critic (usually the same model with a different prompt) evaluates it against the goal, and the agent revises using that feedback. The critique is stored as memory and fed into the next attempt.
GOAL ──▶ ATTEMPT ──▶ CRITIC ──▶ REVISE ──┐
▲ (feedback) │
└──── loop until APPROVED ◀────┘
A minimal reflection loop in code: generate, critique, and only revise while the critic finds problems.
def reflexion(goal, max_revisions=2):
draft = client.messages.create(
model="claude-sonnet-4-6", max_tokens=1024,
messages=[{"role": "user", "content": goal}],
).content[0].text
for _ in range(max_revisions):
critique = client.messages.create(
model="claude-sonnet-4-6", max_tokens=512,
messages=[{"role": "user", "content":
"Goal: " + goal + "\n\nDraft:\n" + draft + "\n\n"
"Critique this draft against the goal. If it fully meets "
"the goal, reply exactly APPROVED. Otherwise list concrete fixes."}],
).content[0].text
if critique.strip().startswith("APPROVED"):
break
draft = client.messages.create(
model="claude-sonnet-4-6", max_tokens=1024,
messages=[{"role": "user", "content":
"Goal: " + goal + "\n\nYour draft:\n" + draft + "\n\n"
"Reviewer feedback:\n" + critique + "\n\nProduce an improved version."}],
).content[0].text
return draft
Reflexion measurably improves quality on tasks with a clear notion of "correct" — code that must pass tests, answers that must satisfy constraints. The cost is obvious in the code: each revision is two extra model calls. You pay roughly 2-3x the tokens for one quality pass.
max_revisions. Two to three rounds capture most of the gain; beyond that you are usually paying for diminishing returns.These patterns are not rivals — they compose. A common production shape is Plan-and-Execute on the outside, ReAct for each step, Reflexion wrapped around steps that must be correct. But for a single starting choice, weigh three axes: latency, cost, and reliability on long tasks.
| Architecture | Latency | Token cost | Best for | Main failure mode |
|---|---|---|---|---|
| ReAct | Low | Low | Short, adaptive tasks; tool-driven Q&A | Wandering, repeated actions on long tasks |
| Plan-and-Execute | Higher (upfront plan) | Medium | Complex multi-step tasks needing structure | Stale plans when reality diverges |
| Reflexion | Highest | High (2-3x) | Quality-critical output with a clear rubric | Oscillating revisions, runaway cost |
A practical decision guide:
Is the task short and tool-driven (a few steps)?
-> ReAct.
Is it long, multi-stage, and you need to audit intent before acting?
-> Plan-and-Execute (with re-planning).
Does correctness matter more than speed, and can you check the output?
-> Wrap the result in Reflexion.
All of the above?
-> Compose them. Plan outside, ReAct per step, Reflexion on critical steps.
Start with the cheapest pattern that could work and add structure only when you observe a failure mode that demands it. Reaching for Plan-and-Execute plus Reflexion on a two-step task buys you latency and a larger bill with no real benefit.
Questions & Answers
stop_reason is no longer "tool_use" (it becomes "end_turn" when Claude is done). Second, the model may be stuck repeating a failing action. Always enforce a max_steps cap, and detect identical consecutive (tool, input) pairs to break early.Key Takeaways
- Every architecture is the same loop. Think → act → observe → repeat, with a terminal state. The architectures differ only in where thinking happens, whether the agent self-critiques, and how it decides to stop.
- ReAct is your default. Interleaved reasoning and acting maps directly onto Claude's tool use API and handles short, adaptive, tool-driven tasks with the least code and cost.
- Plan-and-Execute adds structure. Planning upfront keeps long, multi-stage tasks on track and auditable, at the price of latency — and demands re-planning when observations contradict the plan.
- Reflexion buys quality, not speed. A critique-and-revise loop measurably improves correctness when you can check the output, but each revision costs extra model calls, so cap the rounds.
- Always cap your loops. Uncapped step or revision counts turn a logic bug into a runaway bill.
max_stepsandmax_revisionsare mandatory. - Compose, don't choose blindly. Start cheap, instrument your runs, and add structure only when a real failure mode demands it.
Next Steps: Lesson 3: Tool Use — Giving Agents Hands