Agent Architectures

40 min beginner Lesson 2

Learning Outcomes

  • Diagram the think-act-observe loop that underpins every agent architecture
  • Implement a minimal ReAct loop against Claude's tool use API
  • Contrast Plan-and-Execute with ReAct and explain when each wins
  • Add a Reflexion critique step to improve an agent's output quality
  • Choose an architecture based on latency, cost, and reliability trade-offs

Lesson Plan

Segment Duration Topic
Intro 3 min Why architecture matters more than the model
Explain 6 min The universal think-act-observe loop
Build 8 min Implementing ReAct with Claude tool use
Explain 7 min Plan-and-Execute: separating planning from doing
Build 6 min Adding a Reflexion critique loop
Compare 6 min Trade-offs and a decision table
Wrap-up 4 min Key takeaways and next steps

Before You Begin

Pre-work:

  • Complete Lesson 1: What Are AI Agents? so the planning / tool-use / memory vocabulary is fresh
  • Skim the glossary for the terms agent loop, tool call, and observation
  • Be comfortable making a basic Claude API request in Python

Shopping List:

  • Python 3.10+ with the anthropic SDK installed (pip install anthropic)
  • An Anthropic API key exported as an environment variable
  • A code editor and a terminal
  • A scratch directory for the snippets in this lesson

1 The Universal Loop Behind Every Agent

Before we name any architecture, internalise the one structure they all share. An agent is a model placed inside a loop that lets it act, see the result, and decide again. Strip away the branding and every pattern is a variation on think → act → observe → repeat.

   THINK ──▶ ACT ──▶ OBSERVE ──┐
   (reason)  (tool)  (result)  │
     ▲                         │
     └─── repeat until done ◀──┘
              │
              ▼  FINISH (goal met)

The three things that distinguish architectures are: where the thinking happens (all upfront, or interleaved with acting), whether the agent critiques itself, and how it decides to stop. That last point is the one beginners underestimate — a loop with no stop condition is a bug that bills you per token.

Three terms we will reuse throughout:

Term Meaning
Tool call The model's structured request to run an external function
Observation The result returned to the model after a tool runs
Terminal state The condition that ends the loop (goal met, budget spent, or step cap hit)
NOTE
Key Insight
A chatbot answers in one turn. An agent is the same model wrapped in a loop with tools and a stop condition. The intelligence is in the model; the reliability is in the loop you build around it.
WARNING
Always Cap Your Loops
Every architecture here must enforce a maximum step count. An agent that never recognises it is done can burn your whole budget in minutes. Treat max_steps as non-negotiable.

2 ReAct — Reasoning and Acting, Interleaved

ReAct (Reasoning + Acting) is the simplest and most common architecture. The model alternates between a reasoning thought and a single action, then sees the observation before deciding the next move. There is no separate planning phase — the plan emerges one step at a time.

Here is the control flow in pseudocode:

function react(goal, tools, max_steps):
    messages = [user(goal)]
    for step in range(max_steps):
        response = model.generate(messages, tools)
        if response.is_final_answer:
            return response.text
        for call in response.tool_calls:
            result = run_tool(call.name, call.input)
            messages.append(assistant(call))
            messages.append(tool_result(call.id, result))
    return "Stopped: step limit reached"

The strength is simplicity: the model adapts to whatever the last observation told it. The weakness is that it can wander — without a plan, a ReAct agent on a long task may loop, repeat failed actions, or lose the thread. It is the natural fit for Claude because tool use is the act step: the model requests a tool, you run it, you feed the result back, and the loop continues until the model stops requesting tools.

TIP
Read the Foundational Paper
ReAct comes from Yao et al., 2023, ReAct: Synergizing Reasoning and Acting in Language Models. The core idea is that interleaving a reasoning trace with actions beats either reasoning-only or acting-only on multi-step tasks.
WARNING
Detect Repeated Actions
A ReAct agent can get stuck calling the same tool with the same arguments. Track a short history of (tool name, input) pairs and break the loop or inject a nudge if you see an exact repeat.

3 Implementing ReAct With Claude Tool Use

Let's make it concrete. First, define a tool. Claude's tool use API takes a JSON schema describing each tool's name, purpose, and inputs:

import anthropic

client = anthropic.Anthropic()

TOOLS = [
    {
        "name": "get_word_length",
        "description": "Return the number of characters in a word.",
        "input_schema": {
            "type": "object",
            "properties": {
                "word": {"type": "string", "description": "The word to measure"}
            },
            "required": ["word"],
        },
    }
]

def run_tool(name, tool_input):
    if name == "get_word_length":
        return str(len(tool_input["word"]))
    return f"Unknown tool: {name}"

Now the loop. Notice the three classic moves: send messages, run any requested tools, append the results, repeat. The terminal state is stop_reason == "end_turn" — Claude has stopped asking for tools.

def react_agent(goal, max_steps=8):
    messages = [{"role": "user", "content": goal}]
    for step in range(max_steps):
        resp = client.messages.create(
            model="claude-sonnet-4-6",
            max_tokens=1024,
            tools=TOOLS,
            messages=messages,
        )
        messages.append({"role": "assistant", "content": resp.content})

        if resp.stop_reason != "tool_use":
            return "".join(b.text for b in resp.content if b.type == "text")

        tool_results = []
        for block in resp.content:
            if block.type == "tool_use":
                output = run_tool(block.name, block.input)
                tool_results.append({
                    "type": "tool_result",
                    "tool_use_id": block.id,
                    "content": output,
                })
        messages.append({"role": "user", "content": tool_results})
    return "Stopped: reached max_steps"

print(react_agent("How many letters are in the word 'architecture'?"))

The model reasons ("I should measure the word"), emits a tool_use block, you run it, append a tool_result, and Claude produces the final text answer on the next turn. That is a complete ReAct agent in about thirty lines.

NOTE
Why tool_result Goes in a User Message
In the Anthropic API, tool results are sent back as content blocks inside a user role message. The assistant requested the tool; the user (your harness) supplies the observation. Mismatching the roles is the most common first-time bug.

4 Plan-and-Execute — Decide First, Then Do

ReAct decides one step at a time. Plan-and-Execute splits the work into two phases: a planner produces a full ordered list of steps upfront, then an executor carries out each step (often a small ReAct loop per step). A separate planning pass means the agent commits to a strategy before spending tokens on actions.

   GOAL ──▶ PLANNER ──▶ [step1, step2, step3, ...]
                              │
                              ▼
                          EXECUTOR ──▶ run step, observe ──▶ RE-PLAN?
                              ▲                                 │
                              └──── next step ◀─────────────────┘

You can ask the planner to emit a machine-readable plan so the executor can iterate over it deterministically:

PLAN_PROMPT = """You are a planner. Break the goal into an ordered list of
concrete steps. Respond with ONLY a JSON array of strings, no prose.

Goal: {goal}"""

def make_plan(goal):
    resp = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=512,
        messages=[{"role": "user",
                   "content": PLAN_PROMPT.format(goal=goal)}],
    )
    import json
    return json.loads(resp.content[0].text)

# plan = make_plan("Research the top 3 Python web frameworks and compare them")
# for step in plan:
#     execute_step(step)   # each step can be its own ReAct loop

The big advantage is on complex, multi-step tasks: a plan keeps a long task on track and makes the agent's intent auditable before any side effects occur. The cost is latency and rigidity — you pay for a planning round-trip, and a plan made before seeing reality may be wrong. That is why production Plan-and-Execute systems add a re-plan step: after each executed step, check whether the remaining plan still makes sense, and regenerate it if the world changed.

TIP
Validate the Plan Before Executing
Because the plan is structured JSON, you can sanity-check it cheaply before spending tokens: reject empty plans, cap the number of steps, and confirm each step maps to an available tool. Catching a bad plan here is far cheaper than catching it ten tool calls in.
WARNING
Stale Plans Cause Confident Failures
A plan written upfront assumes the environment matches the model's expectations. If step 2's output contradicts the plan, blindly running steps 3-6 produces confident nonsense. Re-plan when an observation invalidates an assumption.

5 Reflexion — Agents That Critique Themselves

Reflexion adds a self-improvement loop on top of either pattern. The agent produces a draft, a critic (usually the same model with a different prompt) evaluates it against the goal, and the agent revises using that feedback. The critique is stored as memory and fed into the next attempt.

   GOAL ──▶ ATTEMPT ──▶ CRITIC ──▶ REVISE ──┐
              ▲         (feedback)           │
              └──── loop until APPROVED ◀────┘

A minimal reflection loop in code: generate, critique, and only revise while the critic finds problems.

def reflexion(goal, max_revisions=2):
    draft = client.messages.create(
        model="claude-sonnet-4-6", max_tokens=1024,
        messages=[{"role": "user", "content": goal}],
    ).content[0].text

    for _ in range(max_revisions):
        critique = client.messages.create(
            model="claude-sonnet-4-6", max_tokens=512,
            messages=[{"role": "user", "content":
                "Goal: " + goal + "\n\nDraft:\n" + draft + "\n\n"
                "Critique this draft against the goal. If it fully meets "
                "the goal, reply exactly APPROVED. Otherwise list concrete fixes."}],
        ).content[0].text

        if critique.strip().startswith("APPROVED"):
            break

        draft = client.messages.create(
            model="claude-sonnet-4-6", max_tokens=1024,
            messages=[{"role": "user", "content":
                "Goal: " + goal + "\n\nYour draft:\n" + draft + "\n\n"
                "Reviewer feedback:\n" + critique + "\n\nProduce an improved version."}],
        ).content[0].text
    return draft

Reflexion measurably improves quality on tasks with a clear notion of "correct" — code that must pass tests, answers that must satisfy constraints. The cost is obvious in the code: each revision is two extra model calls. You pay roughly 2-3x the tokens for one quality pass.

NOTE
Give the Critic a Concrete Rubric
A vague critic prompt produces vague feedback and endless revisions. Anchor it: paste the failing test output, the acceptance criteria, or a checklist. The more objective the signal, the faster the loop converges and the cleaner the APPROVED stop condition.
WARNING
Cap Revisions or Pay Forever
Self-critique loops can oscillate — fixing one issue reintroduces another. Always set max_revisions. Two to three rounds capture most of the gain; beyond that you are usually paying for diminishing returns.

6 Choosing an Architecture: The Trade-Offs

These patterns are not rivals — they compose. A common production shape is Plan-and-Execute on the outside, ReAct for each step, Reflexion wrapped around steps that must be correct. But for a single starting choice, weigh three axes: latency, cost, and reliability on long tasks.

Architecture Latency Token cost Best for Main failure mode
ReAct Low Low Short, adaptive tasks; tool-driven Q&A Wandering, repeated actions on long tasks
Plan-and-Execute Higher (upfront plan) Medium Complex multi-step tasks needing structure Stale plans when reality diverges
Reflexion Highest High (2-3x) Quality-critical output with a clear rubric Oscillating revisions, runaway cost

A practical decision guide:

Is the task short and tool-driven (a few steps)?
    -> ReAct.

Is it long, multi-stage, and you need to audit intent before acting?
    -> Plan-and-Execute (with re-planning).

Does correctness matter more than speed, and can you check the output?
    -> Wrap the result in Reflexion.

All of the above?
    -> Compose them. Plan outside, ReAct per step, Reflexion on critical steps.

Start with the cheapest pattern that could work and add structure only when you observe a failure mode that demands it. Reaching for Plan-and-Execute plus Reflexion on a two-step task buys you latency and a larger bill with no real benefit.

TIP
Instrument Before You Optimise
Log every step: the tool called, the observation, and a step counter. Most architecture decisions become obvious once you can see where your ReAct agent wanders or how many revisions Reflexion actually needs. Lesson 8 turns this instinct into real evaluation harnesses.

Questions & Answers

Q: My ReAct loop never stops — it keeps calling tools forever. What's wrong?
Two usual causes. First, you may not be checking the right terminal condition — stop when stop_reason is no longer "tool_use" (it becomes "end_turn" when Claude is done). Second, the model may be stuck repeating a failing action. Always enforce a max_steps cap, and detect identical consecutive (tool, input) pairs to break early.
Q: Is Plan-and-Execute always better than ReAct for complex tasks?
No. Plan-and-Execute helps when the task has stable structure you can map out in advance. If the task is highly exploratory — where each step's result fundamentally changes what to do next — a rigid upfront plan goes stale fast, and ReAct's adapt-as-you-go behaviour wins. The middle ground is Plan-and-Execute with re-planning after each step.
Q: Does Reflexion mean I need a second, smarter model to act as the critic?
Not necessarily. The same model with a different, critical prompt is often a perfectly good critic — it is easier to spot flaws in an existing draft than to produce a perfect one in a single pass. A stronger or differently-prompted critic can help on hard tasks, but the bigger lever is giving the critic an objective rubric (tests to pass, constraints to satisfy) rather than asking it to "review for quality."
Q: How do I keep token costs under control across these loops?
Three levers. Cap step and revision counts hard. Keep tool results concise — return only the data the model needs, not entire raw payloads. And prune message history on long ReAct runs; older observations that no longer affect the decision can be summarised or dropped. Memory strategies are the focus of Lesson 4.
Q: Can I mix architectures, or do I have to pick one?
Mix them. These are composable layers, not exclusive choices. The most capable production agents typically plan at the top level, run a ReAct loop to execute each planned step, and apply a Reflexion critique to steps where correctness is non-negotiable. Pick the simplest combination that handles your real failure modes.

Key Takeaways

  1. Every architecture is the same loop. Think → act → observe → repeat, with a terminal state. The architectures differ only in where thinking happens, whether the agent self-critiques, and how it decides to stop.
  2. ReAct is your default. Interleaved reasoning and acting maps directly onto Claude's tool use API and handles short, adaptive, tool-driven tasks with the least code and cost.
  3. Plan-and-Execute adds structure. Planning upfront keeps long, multi-stage tasks on track and auditable, at the price of latency — and demands re-planning when observations contradict the plan.
  4. Reflexion buys quality, not speed. A critique-and-revise loop measurably improves correctness when you can check the output, but each revision costs extra model calls, so cap the rounds.
  5. Always cap your loops. Uncapped step or revision counts turn a logic bug into a runaway bill. max_steps and max_revisions are mandatory.
  6. Compose, don't choose blindly. Start cheap, instrument your runs, and add structure only when a real failure mode demands it.

Next Steps: Lesson 3: Tool Use — Giving Agents Hands