Memory Systems
Learning Outcomes
- Distinguish the four kinds of agent memory and what each one is for
- Manage the short-term context window with summarization and pruning
- Implement long-term memory using a vector store with embeddings
- Design episodic memory that records and replays past agent runs
- Architect a combined memory system for a customer support agent
Lesson Plan
| Segment | Duration | Topic |
|---|---|---|
| Intro | 3 min | Why memory is the missing third leg |
| Concept | 6 min | The four memory types compared |
| Build | 8 min | Short-term: context window management |
| Build | 8 min | Long-term: vector store + embeddings |
| Build | 7 min | Episodic: recording and replaying runs |
| Build | 6 min | Semantic: facts and a knowledge layer |
| Design | 4 min | Combining them for a support agent |
| Wrap-up | 3 min | Key takeaways, preview planning |
Before You Begin
Pre-work:
- Complete Lesson 1: What Are AI Agents? and Lesson 3: Tool Use
- Be comfortable making Anthropic Messages API calls in Python
- Skim the glossary for terms like embedding and context window
Shopping List:
- Python 3.10+ with
anthropicinstalled (pip install anthropic) - An
ANTHROPIC_API_KEYexported in your shell - A vector store — the examples use
chromadb(pip install chromadb), which runs locally with zero setup - A scratch directory for episodic logs (
mkdir agent_memory)
An LLM is stateless. Every API call starts from a blank slate — the model remembers nothing except what you put back into the prompt. "Memory" in agents is the engineering you do around the model to give it continuity. There are four distinct types, and conflating them is the most common architecture mistake.
| Type | Lifespan | Stored in | Retrieved by | Example |
|---|---|---|---|---|
| Short-term | One session | The context window (the message list) | Always present | The current conversation turns |
| Long-term | Across sessions | Vector store / database | Similarity search | "User prefers metric units" |
| Episodic | Across sessions | Append-only event log | Lookup by run/time | "Last Tuesday I solved ticket #441 by restarting the worker" |
| Semantic | Effectively permanent | Knowledge base / facts table | Lookup or RAG | "Refund policy is 30 days" |
The shortcut: short-term is what's in the prompt right now; long-term is what you retrieve in when relevant; episodic is what happened (timestamped events); semantic is what is true (facts independent of any event).
Short-term memory is the running list of messages you send to the model (the standard {"role": ..., "content": ...} turns). It is finite — even large windows fill up over a long agent run, and cost scales with every token you resend. Your job is to keep the window dense with signal. Three techniques bound it:
- Sliding window — keep only the last N turns. Cheap, but you lose early context.
- Summarization — compress older turns into one summary message. Preserves gist, loses detail.
- Token-budget pruning — drop oldest turns until you fit a token ceiling.
Here is summarization-based compaction: once history grows past a threshold, summarize the old turns and replace them, keeping recent turns verbatim.
import anthropic
client = anthropic.Anthropic()
MODEL = "claude-sonnet-4-6" # use your current Sonnet-class model
def compact(messages, keep_recent=4):
"""Summarize all but the most recent turns into one message."""
if len(messages) <= keep_recent + 2:
return messages
old, recent = messages[:-keep_recent], messages[-keep_recent:]
transcript = "\n".join(f"{m['role']}: {m['content']}" for m in old)
summary = client.messages.create(
model=MODEL,
max_tokens=512,
messages=[{"role": "user", "content":
"Summarize this conversation, preserving decisions, open "
"questions, and key facts to remember:\n\n" + transcript}],
)
return [{"role": "user", "content":
f"[Earlier summary]\n{summary.content[0].text}"}] + recent
keep_recent slice) and only compress what is genuinely old.Long-term memory persists across sessions. The dominant pattern: embed each memory as a vector, store it, and retrieve the closest matches by cosine similarity — retrieval-augmented generation (RAG) over the agent's own history. Two phases: write (embed and store) and read (embed the query, search, inject results into the prompt).
import chromadb
# Local, persistent store. No server to run.
store = chromadb.PersistentClient(path="./agent_memory/vectors")
mem = store.get_or_create_collection("long_term")
def remember(text, metadata=None):
"""Write a durable memory. Chroma embeds it for us by default."""
mem.add(ids=[f"mem-{mem.count()}"], documents=[text], metadatas=[metadata or {}])
def recall(query, k=3):
"""Read the k most relevant memories for the current situation."""
results = mem.query(query_texts=[query], n_results=k)
return results["documents"][0] # list of memory strings
Wire it into a turn by retrieving relevant memories and prepending them to the system prompt before answering (Step 6 shows the full assembly). Be selective about what you write: durable signal — stable preferences, resolved problems and their fixes, distilled lessons like "API X rate-limits at N" — and never raw chatter, transient state, or secrets.
Episodic memory answers "what happened, and when?" It is an append-only log of episodes — discrete past experiences with timestamps and outcomes. Where long-term memory holds distilled facts, episodic holds the raw narrative: this goal, these actions, this result. It lets an agent say "last time I hit this error, restarting the worker fixed it." JSON Lines is a durable representation — one object per episode, trivial to append:
import json, time
from pathlib import Path
LOG = Path("./agent_memory/episodes.jsonl")
def record_episode(goal, actions, outcome, success):
episode = {"ts": time.time(), "goal": goal, "actions": actions,
"outcome": outcome, "success": success}
with LOG.open("a") as f:
f.write(json.dumps(episode) + "\n")
def recent_successes(goal_keyword, limit=3):
"""Return past episodes that succeeded on a similar goal, newest first."""
eps = [json.loads(l) for l in LOG.read_text().splitlines()] if LOG.exists() else []
hits = [e for e in eps if e["success"] and goal_keyword.lower() in e["goal"].lower()]
return sorted(hits, key=lambda e: e["ts"], reverse=True)[:limit]
At the start of a run the agent retrieves relevant successful episodes and injects them into the prompt as worked examples ("Previously: deploy 502 -> restarted worker, fixed"), exactly as Step 3 injected long-term memories.
success=false episodes can avoid repeating a dead-end approach — often more valuable than knowing what worked, because the failure space is larger.Semantic memory is the agent's store of facts true independent of any conversation or episode: business rules, product specs, policies, definitions. Unlike episodic memory there is no timestamp narrative — just structured truth the agent looks up.
For small, exact, frequently-changing facts, a structured table beats a vector store — you want deterministic lookups, not fuzzy similarity. Back it with a real data source, then expose it to the agent as a tool so the model fetches authoritative values instead of guessing. This is the Anthropic tool-use schema:
{
"name": "lookup_policy",
"description": "Return an authoritative company policy value. Always use this for refund windows, limits, and supported options instead of relying on memory.",
"input_schema": {
"type": "object",
"properties": {
"key": {
"type": "string",
"enum": ["refund_window_days", "max_file_upload_mb", "support_hours"]
}
},
"required": ["key"]
}
}
The handler is a trivial lookup — KNOWLEDGE.get(key, "UNKNOWN — escalate to a human") against a dict or database row. For larger, unstructured knowledge (docs, manuals, wikis), embed it into the same vector store you use for long-term memory and retrieve via RAG, using a metadata field so a policy lookup never returns a user preference.
lookup_policy tool or RAG so the value is authoritative. The model's job is to reason, not to be the source of truth.Now assemble a memory architecture for a support agent that remembers previous interactions and learns from resolved tickets. Each type plays a distinct role:
| Need | Memory type | Mechanism |
|---|---|---|
| Follow the current chat | Short-term | Message list + compaction |
| Recall this customer's history and prefs | Long-term | Vector store, filtered by customer id |
| Reuse the fix from a similar past ticket | Episodic | Episode log, retrieved by similarity |
| Quote the correct refund policy | Semantic | lookup_policy tool |
Each turn assembles context from all four before calling the model:
def support_turn(customer_id, user_msg, messages):
messages = compact(messages) # short-term
messages.append({"role": "user", "content": user_msg})
prefs = recall(f"customer {customer_id}: {user_msg}", k=3) # long-term
past = recent_successes(goal_keyword=user_msg.split()[0]) # episodic
resolved = [f"{e['goal']} -> {e['outcome']}" for e in past]
system = ("You are a support agent. Use lookup_policy for any policy value.\n"
f"Known about this customer: {prefs}\n"
f"Similar resolved tickets: {resolved}")
resp = client.messages.create(
model=MODEL, max_tokens=1024, system=system,
tools=[POLICY_TOOL], # the lookup_policy schema from Step 5
messages=messages,
)
return resp, messages
When a ticket resolves, write back to memory so the agent improves — call record_episode(...) to log the whole run (episodic) and remember("customer X: resolved '...' via ...", metadata={...}) to distill the durable fact (long-term). This write-back loop is what makes it an agent rather than a stateless chatbot: each resolved ticket makes the next similar one easier.
k values), filter by metadata such as customer id, and measure whether memory actually improves outcomes — see Lesson 8 on evaluation before trusting it in production.Questions & Answers
k results. An empty retrieval beats a confidently irrelevant one.recall should pass that id as a metadata filter so a query can never match another customer's records. Redact secrets and PII at write time — once it's in the index, it's searchable.Key Takeaways
- Memory is four subsystems, not one — short-term (the window), long-term (retrieved facts), episodic (recorded experiences), and semantic (authoritative truth). Design each separately.
- The context window is working memory — keep it dense with sliding windows, summarization, and token budgets; use prompt caching to make resends cheap.
- Long-term memory is RAG over your own history — embed, store, and retrieve top matches, but curate aggressively because a noisy store actively misleads the agent.
- Episodic memory teaches from experience — log episodes with outcomes, retrieve similar past runs, and record failures as deliberately as successes.
- Semantic facts go through tools, not the model's guesses — route policy and limit values through a lookup tool or RAG so answers stay authoritative.
- The write-back loop is what makes it an agent — distilling resolved work back into memory turns yesterday's experience into tomorrow's capability.
Next Steps: Lesson 5: Planning & Reasoning