What Are AI Agents?

30 min beginner Lesson 1

Learning Outcomes

  • Distinguish an agent from a chatbot using a concrete, testable definition
  • Identify the three defining capabilities — tool use, planning, and memory
  • Trace a single iteration of the think → act → observe loop in real code
  • Recognise real-world agents and explain what makes each one agentic
  • Evaluate whether a task needs an agent or just a better prompt

Lesson Plan

Segment Duration Topic
Intro 3 min Why "agent" is the most overloaded word
Define 5 min Chatbot vs agent
Capabilities 12 min The three pillars: tool use, planning, memory
The Loop 5 min Putting it together: a minimal agent loop
Examples 3 min Agents in the wild, and when not to
Wrap-up 2 min Key takeaways and next steps

Before You Begin

Pre-work:

  • This is the first lesson in the Agentic AI course — no prior lesson required.
  • Skim Anthropic's Tool Use guide once; we reference the shapes here.
  • You should have called an LLM API at least once.

Shopping List:

  • Python 3.10+ and pip install anthropic
  • An Anthropic API key exported as ANTHROPIC_API_KEY
  • A terminal and an editor

1 The Testable Definition of an Agent

"Agent" gets slapped on everything from a chained prompt to a fully autonomous coding system. Here's a definition you can apply.

A chatbot is a function: text in, text out. You ask, it answers, the loop ends when you read the reply.

An agent is a system where an LLM decides what action to take next, takes it against the real world, observes the result, and repeats until a goal is met. It isn't producing an answer — it drives a control loop:

Chatbot:  prompt ─> model ─> answer  (done)
Agent:    goal ─> decide ─> run tool ─> observe ─> (goal met? loop)

The test: does the LLM's output change what happens next in the world, and does the world's response feed back into the LLM? If yes, it's an agent. If the LLM just emits text a human reads, it's a chatbot — however clever.

NOTE
Key Insight
Agency is about the control loop, not the model. The same Claude model is a chatbot in one program and an agent in another — the difference is whether you let its output drive actions and feed the results back.

2 Capability One — Tool Use (Giving the Model Hands)

A raw LLM can only generate text. Tool use (function calling) lets the model request that your code run something — a calculation, an API call, a query — and get the result back. You describe each tool with a JSON Schema; the model never runs it, it emits a request your code executes and returns. Here's the shape Claude expects:

weather_tool = {
    "name": "get_weather",
    "description": "Get current weather for a city. Returns temp (C) and conditions.",
    "input_schema": {
        "type": "object",
        "properties": {"city": {"type": "string", "description": "e.g. 'Lisbon'"}},
        "required": ["city"],
    },
}

Ask "What's the weather in Lisbon?" and the model doesn't guess. It returns a structured tool_use block naming get_weather with input {"city": "Lisbon"} and a unique id. Your code calls the real API and returns the result; the model folds it into its answer — the moment it gains hands. Each tool maps to a plain function; more in Lesson 3.

WARNING
The Description Is the Interface
The model selects and parameterises tools entirely from the name and description fields. A vague 'gets data' produces wrong calls — write them as if for a junior engineer who can't see your code: what it does, what it returns, when to use it.

3 Capability Two — Planning (Deciding What to Do)

Tool use alone gives a single action. Planning is the model deciding the sequence of actions needed to reach a goal — and adapting it as it learns.

Take the goal: "Find the cheapest flight from London to Tokyo next month and add it to my calendar." No single tool does this; the model must decompose it — resolve "next month" to dates, call search_flights, compare prices, call create_event.

Two things make this planning rather than a script: decomposition (breaking a fuzzy goal into concrete steps without you spelling them out) and re-planning (if search_flights returns nothing, the model widens the dates and retries — the plan reacts to reality). Elicit it with a system prompt: think step by step, call tools one at a time, observe each result, and reconsider the plan if a step fails.

Planning strategies (interleaved reasoning, plan-then-execute, self-critique) trade off latency, cost, and reliability — covered in Lesson 2 and Lesson 5.

NOTE
Planning Is Why Agents Beat Scripts
If you knew the steps in advance, you'd write a script — faster and cheaper. Agents earn their cost when the path can't be known ahead of time and must be discovered at runtime.

4 Capability Three — Memory (Not Repeating Yourself)

Without memory, an agent forgets every observation the moment it falls out of the context window — re-searching things it found and repeating mistakes. There are two broad kinds, and you'll use both:

Memory Lifespan Backed by
Short-term One task / session The message list in the context window
Long-term Across sessions A database or vector store

Short-term memory is the simplest thing in the world: you keep appending to a list of messages — each tool request and its tool_result — and resend the whole list every turn. That list is the agent's working memory (built up in Step 5). Long-term memory lets an agent improve over time — recalling a past resolution or a user preference — and needs persistent storage: Lesson 4.

WARNING
Context Windows Are Not Memory
A long context window delays the problem; it doesn't solve it. Every token of history costs money and latency, and quality degrades as the window fills. Treat it as scratch space.

5 The Agent Loop — Putting It Together

The three capabilities combine into one repeating cycle: think → act → observe. Here is a minimal loop using Claude's tool use API:

import anthropic
client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from env

def get_weather(city: str) -> str:
    return f"{city}: 19C, partly cloudy"  # stub; call a real API in practice

TOOLS = [weather_tool]  # the schema from Step 2
messages = [{"role": "user", "content": "Should I take a coat in Oslo today?"}]

while True:
    resp = client.messages.create(
        model="claude-sonnet-4-6", max_tokens=1024, tools=TOOLS, messages=messages)
    messages.append({"role": "assistant", "content": resp.content})

    if resp.stop_reason != "tool_use":      # model has its final answer
        print(resp.content[-1].text)
        break

    results = []                            # run every tool the model asked for
    for b in resp.content:
        if b.type == "tool_use":
            out = get_weather(**b.input)
            results.append({"type": "tool_result",
                            "tool_use_id": b.id, "content": out})
    messages.append({"role": "user", "content": results})

One iteration: think (the model decides it needs the weather, so stop_reason is tool_use), act (your code runs get_weather), observe (you append a tool_result and loop). Next pass the model answers in plain text and exits. The model only decides "call a tool" or "done" — the loop and stopping logic are all your code.

WARNING
Always Bound the Loop
A bare while True can spin forever or rack up cost. In production you cap iterations (e.g. for _ in range(10)), set a token budget, and add a timeout — see Lesson 7.

6 Agents in the Wild — and When NOT to Use One

You've used agents already, maybe without naming them:

Agent Goal What makes it agentic
Coding assistant "Fix this failing test" Edits, runs tests, reads failures, edits again until green
Research agent "Summarise the state of X" Searches, reads, decides what to search next
Data agent "Why did signups drop?" Queries, inspects results, forms a new query

The common thread: a goal, tools, and a loop that runs until the goal is met or a limit is hit. But agents are slower, costlier, and less predictable than a direct call. Reach for a plain LLM call — or no LLM at all — when the steps are fixed (write a script: faster, cheaper, deterministic), when a single prompt suffices ("summarise this" needs no tools or loop), or when a wrong action is costly and you have no guardrails (production write access needs Lesson 7 first).

NOTE
The Honest Rule of Thumb
Use an agent when the path to the goal must be discovered at runtime and varies per request. If the path is the same every time, you want a workflow — many 'agent' projects are really workflows that would be more reliable as code.

Questions & Answers

Q: Isn't this just function calling with a loop? Why the new vocabulary?
Mechanically, yes. The framing matters because it shifts your focus to the hard parts: when does the loop stop, how do you bound cost, how does the model recover from a failed tool, how do you keep state. None of those exist for a single call — and they're where agents break.
Q: What stops the agent from looping forever or burning my whole budget?
Nothing, unless you add it. The Step 5 loop is unbounded for clarity. In production you cap iterations, set a token budget, add a timeout, and often require approval before destructive actions — see Lesson 7.
Q: Do I need a framework like LangChain or LangGraph to build agents?
No. As Step 5 shows, the loop is a few lines around the SDK. Frameworks help with complex graphs, persistence, and observability, but the raw loop teaches you what they're doing. Learn the loop first; add a framework when it earns its keep.
Q: How is an agent different from RAG (retrieval-augmented generation)?
RAG retrieves documents into the prompt before a single generation — a one-shot pipeline. An agent can decide mid-task to retrieve, act on what it found, then retrieve again with a refined query. RAG is a step an agent might take; alone it has no loop.
Q: Can the model lie about a tool result or hallucinate that it called one?
Yes, especially if you let it write free-text claims like "I've sent the email." That's why the loop trusts only structured tool_use blocks and real tool_result values. Never let the model's prose be the source of truth for whether an action happened; your code's execution record is.

Key Takeaways

  1. An agent is a control loop, not a model. The test: does the LLM's output drive real actions, and do the results feed back in?
  2. Three capabilities define agents. Tool use (hands), planning (the next step), and memory (not repeating yourself) — remove one and it's less than an agent.
  3. The loop is think → act → observe. A few lines around an API call. You own the loop; the model only chooses "call a tool" or "done."
  4. Context windows are not memory. Short-term state lives in the message list; durable memory needs storage. A bigger window only delays the problem.
  5. Agents are not always the answer. If the steps are fixed, write a workflow. Reach for an agent only when the path must be discovered at runtime.

Next Steps: Lesson 2: Agent Architectures