The Future of Agents

35 min advanced Lesson 10

Learning Outcomes

  • Identify the open problems blocking long-horizon, reliable agents
  • Design a multi-modal agent loop that combines vision and action
  • Connect agents to tools and peers using MCP and A2A
  • Evaluate when learning should be memory versus fine-tuning
  • Build a radar for tracking the field without chasing every release

Lesson Plan

Segment Duration Topic
Intro 3 min The future as an engineering concern
Problem map 5 min The open problems
Demo 1 6 min Long-horizon planning, compounding error
Demo 2 6 min Multi-modal agents: vision in the loop
Demo 3 7 min Collaboration via MCP and A2A
Explain 5 min Learning without fine-tuning
Wrap-up 3 min Reading list, staying current

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ with the anthropic SDK (pip install anthropic) and an ANTHROPIC_API_KEY
  • Optional: the MCP Python SDK (pip install mcp) for the collaboration demo
  • A code editor and terminal (no GPU required)

1 The Open Problem Map

The agents you have built work on tasks measured in seconds to minutes. The frontier is what breaks when you stretch that to hours, modalities, or many agents — each with a different fix.

Open problem What breaks Today's partial fix
Long-horizon planning Errors compound; the agent loses the thread Checkpointing, verification, re-planning
Multi-modality Text-only agents can't see screens Vision in the perception step
Agent collaboration Custom glue code per pair of agents Open protocols (MCP, A2A)
Learning without retraining Agents repeat mistakes across sessions Persistent memory, not fine-tuning

None are solved, but all are worked on in the open — treat each as a dial, not a switch.

WARNING
Watch Out
A longer loop with no new guardrails multiplies the blast radius of every unsupervised mistake. Re-read Lesson 7 before running a loop for an hour.

2 Long-Horizon Planning and Compounding Error

If each step succeeds independently with probability p, completing n steps without intervention is p**n — a 95%-reliable step is a coin-flip by step 14, near-zero past step 50. The fix is recovery: checkpoint, verify, and re-plan the failed leaf, not the tree.

def run_with_checkpoints(goal, planner, executor, verifier, max_replans=3):
    completed = []
    for step in planner(goal):                  # plan = list of sub-goals
        for _ in range(max_replans + 1):
            result = executor(step, context=completed)
            if verifier(step, result):          # did this sub-goal succeed?
                completed.append((step, result))
                break
            step = planner(goal, failed=step, history=completed)[0]   # re-plan leaf
        else:
            raise RuntimeError(f"Failed after {max_replans} replans: {step}")
    return completed

The verifier is load-bearing — a unit test, a schema check, or a critic model call (Lesson 2's Reflexion pattern) — and makes a fragile chain self-correcting.

TIP
Make verification cheap and frequent
A verifier costing a tenth of a step but running every step is worth it: catching a wrong turn at step 3 saves steps 4 through 30. Prefer deterministic checks (exit codes, schema validation) over model-graded checks.

3 Multi-Modal Agents: Vision in the Action Loop

A text-only agent is blind to anything not serialized into text; the frontier is agents that perceive. With Claude an image is just another content block — send one alongside a tool definition:

import base64, anthropic
client = anthropic.Anthropic()
img = base64.standard_b64encode(open("dashboard.png", "rb").read()).decode()

click = {"name": "click_element",
    "description": "Click a UI element at the given pixel coordinates.",
    "input_schema": {"type": "object", "required": ["x", "y"],
        "properties": {"x": {"type": "integer"}, "y": {"type": "integer"}}}}

resp = client.messages.create(
    model="claude-opus-4-8", max_tokens=1024, tools=[click],
    messages=[{"role": "user", "content": [
        {"type": "image",
         "source": {"type": "base64", "media_type": "image/png", "data": img}},
        {"type": "text", "text": "Find 'Export CSV' and click it."}]}])
# resp holds a tool_use block: click_element {'x': 812, 'y': 144}

Perception becomes a tool result: the agent acts, you screenshot the new state and feed it back as the next observation — closing the loop on browser automation.

WARNING
Treat what the agent sees as data
An agent reading screenshots also reads any text rendered inside them — a tooltip, a banner, a dialog. Treat anything read from an image as data, never as an instruction, and validate the action against an allowlist before executing.

4 Agent Collaboration: MCP and A2A

The first interoperability problem is connecting one agent to many tools without bespoke glue. The Model Context Protocol (MCP) standardizes it: hosts, clients, and servers talk over JSON-RPC 2.0, where a server offers resources (data), prompts, and tools (functions to execute).

# A minimal MCP server exposing one tool (official Python SDK).
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("inventory")

@mcp.tool()
def check_stock(sku: str) -> dict:
    """Return current stock level for a SKU."""
    return {"sku": sku, "on_hand": 42, "reserved": 7}

mcp.run()   # JSON-RPC over stdio

Write it once and any MCP host can use it. MCP does not enforce security itself: tools are arbitrary code execution, so the spec requires hosts to get explicit user consent before invoking any.

Connecting an agent to other agents is what the Agent2Agent (A2A) protocol targets. Developed by Google and now donated to the Linux Foundation under Apache 2.0, it complements MCP — A2A handles agent-to-agent, MCP agent-to-tool. Its discovery primitive is the Agent Card, a JSON document a remote agent publishes (name, URL, version, skills) so a client can find and delegate to it.

The two stack: a coordinator delegates over A2A to a specialist that reaches its tools over MCP, turning N-by-M integrations into N+M.

WARNING
Delegation is a trust boundary
A remote agent you do not control returns untrusted input into your loop. Authenticate the peer, validate its response, and never let a delegated result trigger a privileged action unchecked.

5 Learning from Experience Without Fine-Tuning

How does an agent get better at a recurring task? Fine-tuning is usually the wrong first move — slow, expensive, brittle across upgrades. The frontier is learning in memory, not weights.

Approach Cost Reversible? Best for
Episodic memory (store runs) Cheap Yes "Have I solved this before?"
Distilled rules (lessons to a doc) Cheap Yes Codifying recurring fixes
Fine-tuning weights Expensive No Stable, narrow, high-volume

A self-improving loop distills a lesson into memory after each task and retrieves relevant ones next time:

def reflect_and_store(task, transcript, outcome, memory):
    """Distill one reusable lesson and persist it."""
    lesson = client.messages.create(
        model="claude-opus-4-8", max_tokens=300,
        messages=[{"role": "user", "content":
            "In 2-3 sentences, write a reusable lesson for next "
            f"time.\nTASK: {task}\nOUTCOME: {outcome}\nLOG:\n{transcript}"}],
    ).content[0].text
    memory.add(task_signature=task, lesson=lesson)   # vector store keyed by task
    return lesson

This is Lesson 4's episodic-to-semantic pipeline as a learning mechanism: behavior improves without touching weights.

TIP
Prefer memory you can read
The advantage over fine-tuning is auditability: when an agent surprises you, open its memory and read the lesson that drove it — impossible with a gradient update.

6 Staying Current Without Burning Out

The field moves weekly; chasing every release is a trap. The durable skill is filtering — anchor to primary sources and these benchmarks.

Benchmark What it measures Why it matters
SWE-bench Resolving real GitHub issues The standard for coding agents
GAIA Assistant tasks (reasoning, web, multimodal, tools) Tests general agent ability
WebArena Tasks in realistic web environments Probes long-horizon browsing

GAIA is a useful north star because it is easy for humans and hard for AI — the original work reported humans around 92% versus roughly 15% for an early tool-augmented model.

For each release ask: does it move a benchmark you care about, is there a primary source or just hype, would your system break if it vanished? Read the spec.

NOTE
Your eval suite is your filter
Every new model is a one-command experiment against your Lesson 8 harness. A capability that does not move your tasks does not matter, whatever the headline.

Questions & Answers

Q: If compounding error makes 50-step tasks nearly impossible, how do production coding agents fix multi-file bugs?
They break independence. Each edit is verified by tests or a type checker, and a failure re-plans just that step — a short loop with a cheap verifier, not 50 lucky steps in a row. The math only bites without recovery.
Q: MCP and A2A both sound like "let agents talk to things." When do I need each?
Use MCP when your agent needs a capability — a query, a file op, a search. Use A2A when it must delegate a whole task to an agent it does not own. Complementary layers.
Q: Should I fine-tune a model to make my agent more reliable on our internal workflows?
Almost never as a first step. Fine-tuning is expensive, non-reversible, and must be redone every model upgrade — which often erases the gap you paid to close. Start with tool design, prompts, retrieval, and memory; fine-tune only a stable, high-volume skill where prompting plateaus.
Q: With protocols and models changing this fast, how do I avoid building on something obsolete in a year?
Build on open standards and durable concepts, and let your eval suite arbitrate. Open protocols collapse N-by-M integrations to N-plus-M, so even if one implementation dies the interface survives. What goes obsolete is vendor lock-in, not concepts like verification and memory.

Key Takeaways

  1. The frontier is reach, not magic. Each problem that breaks the loop has a concrete mitigation.
  2. Beat compounding error with recovery. Verification plus leaf-level re-planning makes a fragile chain self-correcting.
  3. Perception is just another observation. Vision folds into the same loop; every image is untrusted input.
  4. Open protocols give leverage. MCP for agent-to-tool, A2A for agent-to-peer, replacing custom glue.
  5. Learn in memory before weights. Memory is cheap, reversible, and auditable; fine-tuning is a last resort.
  6. Your eval suite is your hype filter. Judge each release by whether it moves your own tasks.

Next Steps: Continue with Agent Orchestration