What Are AI Agents?
Learning Outcomes
- Distinguish an agent from a chatbot using a concrete, testable definition
- Identify the three defining capabilities — tool use, planning, and memory
- Trace a single iteration of the think → act → observe loop in real code
- Recognise real-world agents and explain what makes each one agentic
- Evaluate whether a task needs an agent or just a better prompt
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why "agent" is the most overloaded word |
| Define | 5 min | Chatbot vs agent |
| Capabilities | 12 min | The three pillars: tool use, planning, memory |
| The Loop | 5 min | Putting it together: a minimal agent loop |
| Examples | 3 min | Agents in the wild, and when not to |
| Wrap-up | 2 min | Key takeaways and next steps |
Before You Begin
Pre-work:
- This is the first lesson in the Agentic AI course — no prior lesson required.
- Skim Anthropic's Tool Use guide once; we reference the shapes here.
- You should have called an LLM API at least once.
Shopping List:
- Python 3.10+ and
pip install anthropic - An Anthropic API key exported as
ANTHROPIC_API_KEY - A terminal and an editor
"Agent" gets slapped on everything from a chained prompt to a fully autonomous coding system. Here's a definition you can apply.
A chatbot is a function: text in, text out. You ask, it answers, the loop ends when you read the reply.
An agent is a system where an LLM decides what action to take next, takes it against the real world, observes the result, and repeats until a goal is met. It isn't producing an answer — it drives a control loop:
Chatbot: prompt ─> model ─> answer (done)
Agent: goal ─> decide ─> run tool ─> observe ─> (goal met? loop)
The test: does the LLM's output change what happens next in the world, and does the world's response feed back into the LLM? If yes, it's an agent. If the LLM just emits text a human reads, it's a chatbot — however clever.
A raw LLM can only generate text. Tool use (function calling) lets the model request that your code run something — a calculation, an API call, a query — and get the result back. You describe each tool with a JSON Schema; the model never runs it, it emits a request your code executes and returns. Here's the shape Claude expects:
weather_tool = {
"name": "get_weather",
"description": "Get current weather for a city. Returns temp (C) and conditions.",
"input_schema": {
"type": "object",
"properties": {"city": {"type": "string", "description": "e.g. 'Lisbon'"}},
"required": ["city"],
},
}
Ask "What's the weather in Lisbon?" and the model doesn't guess. It returns a structured tool_use block naming get_weather with input {"city": "Lisbon"} and a unique id. Your code calls the real API and returns the result; the model folds it into its answer — the moment it gains hands. Each tool maps to a plain function; more in Lesson 3.
name and description fields. A vague 'gets data' produces wrong calls — write them as if for a junior engineer who can't see your code: what it does, what it returns, when to use it.Tool use alone gives a single action. Planning is the model deciding the sequence of actions needed to reach a goal — and adapting it as it learns.
Take the goal: "Find the cheapest flight from London to Tokyo next month and add it to my calendar." No single tool does this; the model must decompose it — resolve "next month" to dates, call search_flights, compare prices, call create_event.
Two things make this planning rather than a script: decomposition (breaking a fuzzy goal into concrete steps without you spelling them out) and re-planning (if search_flights returns nothing, the model widens the dates and retries — the plan reacts to reality). Elicit it with a system prompt: think step by step, call tools one at a time, observe each result, and reconsider the plan if a step fails.
Planning strategies (interleaved reasoning, plan-then-execute, self-critique) trade off latency, cost, and reliability — covered in Lesson 2 and Lesson 5.
Without memory, an agent forgets every observation the moment it falls out of the context window — re-searching things it found and repeating mistakes. There are two broad kinds, and you'll use both:
| Memory | Lifespan | Backed by |
|---|---|---|
| Short-term | One task / session | The message list in the context window |
| Long-term | Across sessions | A database or vector store |
Short-term memory is the simplest thing in the world: you keep appending to a list of messages — each tool request and its tool_result — and resend the whole list every turn. That list is the agent's working memory (built up in Step 5). Long-term memory lets an agent improve over time — recalling a past resolution or a user preference — and needs persistent storage: Lesson 4.
The three capabilities combine into one repeating cycle: think → act → observe. Here is a minimal loop using Claude's tool use API:
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from env
def get_weather(city: str) -> str:
return f"{city}: 19C, partly cloudy" # stub; call a real API in practice
TOOLS = [weather_tool] # the schema from Step 2
messages = [{"role": "user", "content": "Should I take a coat in Oslo today?"}]
while True:
resp = client.messages.create(
model="claude-sonnet-4-6", max_tokens=1024, tools=TOOLS, messages=messages)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use": # model has its final answer
print(resp.content[-1].text)
break
results = [] # run every tool the model asked for
for b in resp.content:
if b.type == "tool_use":
out = get_weather(**b.input)
results.append({"type": "tool_result",
"tool_use_id": b.id, "content": out})
messages.append({"role": "user", "content": results})
One iteration: think (the model decides it needs the weather, so stop_reason is tool_use), act (your code runs get_weather), observe (you append a tool_result and loop). Next pass the model answers in plain text and exits. The model only decides "call a tool" or "done" — the loop and stopping logic are all your code.
while True can spin forever or rack up cost. In production you cap iterations (e.g. for _ in range(10)), set a token budget, and add a timeout — see Lesson 7.You've used agents already, maybe without naming them:
| Agent | Goal | What makes it agentic |
|---|---|---|
| Coding assistant | "Fix this failing test" | Edits, runs tests, reads failures, edits again until green |
| Research agent | "Summarise the state of X" | Searches, reads, decides what to search next |
| Data agent | "Why did signups drop?" | Queries, inspects results, forms a new query |
The common thread: a goal, tools, and a loop that runs until the goal is met or a limit is hit. But agents are slower, costlier, and less predictable than a direct call. Reach for a plain LLM call — or no LLM at all — when the steps are fixed (write a script: faster, cheaper, deterministic), when a single prompt suffices ("summarise this" needs no tools or loop), or when a wrong action is costly and you have no guardrails (production write access needs Lesson 7 first).
Questions & Answers
tool_use blocks and real tool_result values. Never let the model's prose be the source of truth for whether an action happened; your code's execution record is.Key Takeaways
- An agent is a control loop, not a model. The test: does the LLM's output drive real actions, and do the results feed back in?
- Three capabilities define agents. Tool use (hands), planning (the next step), and memory (not repeating yourself) — remove one and it's less than an agent.
- The loop is think → act → observe. A few lines around an API call. You own the loop; the model only chooses "call a tool" or "done."
- Context windows are not memory. Short-term state lives in the message list; durable memory needs storage. A bigger window only delays the problem.
- Agents are not always the answer. If the steps are fixed, write a workflow. Reach for an agent only when the path must be discovered at runtime.
Next Steps: Lesson 2: Agent Architectures