The Future of Agents
Learning Outcomes
- Identify the open problems blocking long-horizon, reliable agents
- Design a multi-modal agent loop that combines vision and action
- Connect agents to tools and peers using MCP and A2A
- Evaluate when learning should be memory versus fine-tuning
- Build a radar for tracking the field without chasing every release
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | The future as an engineering concern |
| Problem map | 5 min | The open problems |
| Demo 1 | 6 min | Long-horizon planning, compounding error |
| Demo 2 | 6 min | Multi-modal agents: vision in the loop |
| Demo 3 | 7 min | Collaboration via MCP and A2A |
| Explain | 5 min | Learning without fine-tuning |
| Wrap-up | 3 min | Reading list, staying current |
Before You Begin
Pre-work:
- Complete Lesson 6 — Building Your First Agent so you have a loop to extend
- Review Lesson 7 — Agent Safety & Guardrails and Lesson 8 — Evaluating Agent Performance — every direction here multiplies risk
Shopping List:
- Python 3.10+ with the
anthropicSDK (pip install anthropic) and anANTHROPIC_API_KEY - Optional: the MCP Python SDK (
pip install mcp) for the collaboration demo - A code editor and terminal (no GPU required)
The agents you have built work on tasks measured in seconds to minutes. The frontier is what breaks when you stretch that to hours, modalities, or many agents — each with a different fix.
| Open problem | What breaks | Today's partial fix |
|---|---|---|
| Long-horizon planning | Errors compound; the agent loses the thread | Checkpointing, verification, re-planning |
| Multi-modality | Text-only agents can't see screens | Vision in the perception step |
| Agent collaboration | Custom glue code per pair of agents | Open protocols (MCP, A2A) |
| Learning without retraining | Agents repeat mistakes across sessions | Persistent memory, not fine-tuning |
None are solved, but all are worked on in the open — treat each as a dial, not a switch.
If each step succeeds independently with probability p, completing n steps without intervention is p**n — a 95%-reliable step is a coin-flip by step 14, near-zero past step 50. The fix is recovery: checkpoint, verify, and re-plan the failed leaf, not the tree.
def run_with_checkpoints(goal, planner, executor, verifier, max_replans=3):
completed = []
for step in planner(goal): # plan = list of sub-goals
for _ in range(max_replans + 1):
result = executor(step, context=completed)
if verifier(step, result): # did this sub-goal succeed?
completed.append((step, result))
break
step = planner(goal, failed=step, history=completed)[0] # re-plan leaf
else:
raise RuntimeError(f"Failed after {max_replans} replans: {step}")
return completed
The verifier is load-bearing — a unit test, a schema check, or a critic model call (Lesson 2's Reflexion pattern) — and makes a fragile chain self-correcting.
A text-only agent is blind to anything not serialized into text; the frontier is agents that perceive. With Claude an image is just another content block — send one alongside a tool definition:
import base64, anthropic
client = anthropic.Anthropic()
img = base64.standard_b64encode(open("dashboard.png", "rb").read()).decode()
click = {"name": "click_element",
"description": "Click a UI element at the given pixel coordinates.",
"input_schema": {"type": "object", "required": ["x", "y"],
"properties": {"x": {"type": "integer"}, "y": {"type": "integer"}}}}
resp = client.messages.create(
model="claude-opus-4-8", max_tokens=1024, tools=[click],
messages=[{"role": "user", "content": [
{"type": "image",
"source": {"type": "base64", "media_type": "image/png", "data": img}},
{"type": "text", "text": "Find 'Export CSV' and click it."}]}])
# resp holds a tool_use block: click_element {'x': 812, 'y': 144}
Perception becomes a tool result: the agent acts, you screenshot the new state and feed it back as the next observation — closing the loop on browser automation.
The first interoperability problem is connecting one agent to many tools without bespoke glue. The Model Context Protocol (MCP) standardizes it: hosts, clients, and servers talk over JSON-RPC 2.0, where a server offers resources (data), prompts, and tools (functions to execute).
# A minimal MCP server exposing one tool (official Python SDK).
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("inventory")
@mcp.tool()
def check_stock(sku: str) -> dict:
"""Return current stock level for a SKU."""
return {"sku": sku, "on_hand": 42, "reserved": 7}
mcp.run() # JSON-RPC over stdio
Write it once and any MCP host can use it. MCP does not enforce security itself: tools are arbitrary code execution, so the spec requires hosts to get explicit user consent before invoking any.
Connecting an agent to other agents is what the Agent2Agent (A2A) protocol targets. Developed by Google and now donated to the Linux Foundation under Apache 2.0, it complements MCP — A2A handles agent-to-agent, MCP agent-to-tool. Its discovery primitive is the Agent Card, a JSON document a remote agent publishes (name, URL, version, skills) so a client can find and delegate to it.
The two stack: a coordinator delegates over A2A to a specialist that reaches its tools over MCP, turning N-by-M integrations into N+M.
How does an agent get better at a recurring task? Fine-tuning is usually the wrong first move — slow, expensive, brittle across upgrades. The frontier is learning in memory, not weights.
| Approach | Cost | Reversible? | Best for |
|---|---|---|---|
| Episodic memory (store runs) | Cheap | Yes | "Have I solved this before?" |
| Distilled rules (lessons to a doc) | Cheap | Yes | Codifying recurring fixes |
| Fine-tuning weights | Expensive | No | Stable, narrow, high-volume |
A self-improving loop distills a lesson into memory after each task and retrieves relevant ones next time:
def reflect_and_store(task, transcript, outcome, memory):
"""Distill one reusable lesson and persist it."""
lesson = client.messages.create(
model="claude-opus-4-8", max_tokens=300,
messages=[{"role": "user", "content":
"In 2-3 sentences, write a reusable lesson for next "
f"time.\nTASK: {task}\nOUTCOME: {outcome}\nLOG:\n{transcript}"}],
).content[0].text
memory.add(task_signature=task, lesson=lesson) # vector store keyed by task
return lesson
This is Lesson 4's episodic-to-semantic pipeline as a learning mechanism: behavior improves without touching weights.
The field moves weekly; chasing every release is a trap. The durable skill is filtering — anchor to primary sources and these benchmarks.
| Benchmark | What it measures | Why it matters |
|---|---|---|
| SWE-bench | Resolving real GitHub issues | The standard for coding agents |
| GAIA | Assistant tasks (reasoning, web, multimodal, tools) | Tests general agent ability |
| WebArena | Tasks in realistic web environments | Probes long-horizon browsing |
GAIA is a useful north star because it is easy for humans and hard for AI — the original work reported humans around 92% versus roughly 15% for an early tool-augmented model.
For each release ask: does it move a benchmark you care about, is there a primary source or just hype, would your system break if it vanished? Read the spec.
Questions & Answers
Key Takeaways
- The frontier is reach, not magic. Each problem that breaks the loop has a concrete mitigation.
- Beat compounding error with recovery. Verification plus leaf-level re-planning makes a fragile chain self-correcting.
- Perception is just another observation. Vision folds into the same loop; every image is untrusted input.
- Open protocols give leverage. MCP for agent-to-tool, A2A for agent-to-peer, replacing custom glue.
- Learn in memory before weights. Memory is cheap, reversible, and auditable; fine-tuning is a last resort.
- Your eval suite is your hype filter. Judge each release by whether it moves your own tasks.
Next Steps: Continue with Agent Orchestration