Memory Systems

45 min intermediate Lesson 4

Learning Outcomes

  • Distinguish the four kinds of agent memory and what each one is for
  • Manage the short-term context window with summarization and pruning
  • Implement long-term memory using a vector store with embeddings
  • Design episodic memory that records and replays past agent runs
  • Architect a combined memory system for a customer support agent

Lesson Plan

Segment Duration Topic
Intro 3 min Why memory is the missing third leg
Concept 6 min The four memory types compared
Build 8 min Short-term: context window management
Build 8 min Long-term: vector store + embeddings
Build 7 min Episodic: recording and replaying runs
Build 6 min Semantic: facts and a knowledge layer
Design 4 min Combining them for a support agent
Wrap-up 3 min Key takeaways, preview planning

Before You Begin

Pre-work:

Shopping List:

  • Python 3.10+ with anthropic installed (pip install anthropic)
  • An ANTHROPIC_API_KEY exported in your shell
  • A vector store — the examples use chromadb (pip install chromadb), which runs locally with zero setup
  • A scratch directory for episodic logs (mkdir agent_memory)

1 The Four Memory Types

An LLM is stateless. Every API call starts from a blank slate — the model remembers nothing except what you put back into the prompt. "Memory" in agents is the engineering you do around the model to give it continuity. There are four distinct types, and conflating them is the most common architecture mistake.

Type Lifespan Stored in Retrieved by Example
Short-term One session The context window (the message list) Always present The current conversation turns
Long-term Across sessions Vector store / database Similarity search "User prefers metric units"
Episodic Across sessions Append-only event log Lookup by run/time "Last Tuesday I solved ticket #441 by restarting the worker"
Semantic Effectively permanent Knowledge base / facts table Lookup or RAG "Refund policy is 30 days"

The shortcut: short-term is what's in the prompt right now; long-term is what you retrieve in when relevant; episodic is what happened (timestamped events); semantic is what is true (facts independent of any event).

NOTE
Key Insight
Memory is not one system. Short-term is bounded by the context window and managed by truncation; the other three are external stores you query and selectively inject. Treat them as separate subsystems with separate retrieval strategies.
WARNING
Watch Out
Do not make the context window your only memory. Stuffing the entire history into every call is expensive, slow, and degrades quality as the window fills with low-signal tokens. The window is working memory, not storage.

2 Short-Term Memory: Managing the Context Window

Short-term memory is the running list of messages you send to the model (the standard {"role": ..., "content": ...} turns). It is finite — even large windows fill up over a long agent run, and cost scales with every token you resend. Your job is to keep the window dense with signal. Three techniques bound it:

  1. Sliding window — keep only the last N turns. Cheap, but you lose early context.
  2. Summarization — compress older turns into one summary message. Preserves gist, loses detail.
  3. Token-budget pruning — drop oldest turns until you fit a token ceiling.

Here is summarization-based compaction: once history grows past a threshold, summarize the old turns and replace them, keeping recent turns verbatim.

import anthropic

client = anthropic.Anthropic()
MODEL = "claude-sonnet-4-6"  # use your current Sonnet-class model

def compact(messages, keep_recent=4):
    """Summarize all but the most recent turns into one message."""
    if len(messages) <= keep_recent + 2:
        return messages
    old, recent = messages[:-keep_recent], messages[-keep_recent:]
    transcript = "\n".join(f"{m['role']}: {m['content']}" for m in old)
    summary = client.messages.create(
        model=MODEL,
        max_tokens=512,
        messages=[{"role": "user", "content":
            "Summarize this conversation, preserving decisions, open "
            "questions, and key facts to remember:\n\n" + transcript}],
    )
    return [{"role": "user", "content":
             f"[Earlier summary]\n{summary.content[0].text}"}] + recent
TIP
Tip
If you call the Anthropic API directly, turn on prompt caching for the stable prefix of your message list (system prompt, tool definitions, early turns). You pay full price once, then a fraction on cache hits — making resends far cheaper with no code restructuring.
WARNING
Watch Out
Summarization is lossy and silent. A summary can drop the one detail that mattered. Keep the most recent turns verbatim (the keep_recent slice) and only compress what is genuinely old.

3 Long-Term Memory: A Vector Store

Long-term memory persists across sessions. The dominant pattern: embed each memory as a vector, store it, and retrieve the closest matches by cosine similarity — retrieval-augmented generation (RAG) over the agent's own history. Two phases: write (embed and store) and read (embed the query, search, inject results into the prompt).

import chromadb

# Local, persistent store. No server to run.
store = chromadb.PersistentClient(path="./agent_memory/vectors")
mem = store.get_or_create_collection("long_term")

def remember(text, metadata=None):
    """Write a durable memory. Chroma embeds it for us by default."""
    mem.add(ids=[f"mem-{mem.count()}"], documents=[text], metadatas=[metadata or {}])

def recall(query, k=3):
    """Read the k most relevant memories for the current situation."""
    results = mem.query(query_texts=[query], n_results=k)
    return results["documents"][0]  # list of memory strings

Wire it into a turn by retrieving relevant memories and prepending them to the system prompt before answering (Step 6 shows the full assembly). Be selective about what you write: durable signal — stable preferences, resolved problems and their fixes, distilled lessons like "API X rate-limits at N" — and never raw chatter, transient state, or secrets.

NOTE
Key Insight
Retrieval quality is dominated by what you store, not by the LLM. A noisy store returns irrelevant memories that actively mislead the agent. Curate writes: a small store of high-signal facts beats a huge store of raw logs.
WARNING
Watch Out
Embeddings encode whatever you put in them, including PII and secrets. Apply the same redaction policy to memory writes that you would to logs — never persist credentials or tokens to a vector store.

4 Episodic Memory: Recording and Replaying Runs

Episodic memory answers "what happened, and when?" It is an append-only log of episodes — discrete past experiences with timestamps and outcomes. Where long-term memory holds distilled facts, episodic holds the raw narrative: this goal, these actions, this result. It lets an agent say "last time I hit this error, restarting the worker fixed it." JSON Lines is a durable representation — one object per episode, trivial to append:

import json, time
from pathlib import Path

LOG = Path("./agent_memory/episodes.jsonl")

def record_episode(goal, actions, outcome, success):
    episode = {"ts": time.time(), "goal": goal, "actions": actions,
               "outcome": outcome, "success": success}
    with LOG.open("a") as f:
        f.write(json.dumps(episode) + "\n")

def recent_successes(goal_keyword, limit=3):
    """Return past episodes that succeeded on a similar goal, newest first."""
    eps = [json.loads(l) for l in LOG.read_text().splitlines()] if LOG.exists() else []
    hits = [e for e in eps if e["success"] and goal_keyword.lower() in e["goal"].lower()]
    return sorted(hits, key=lambda e: e["ts"], reverse=True)[:limit]

At the start of a run the agent retrieves relevant successful episodes and injects them into the prompt as worked examples ("Previously: deploy 502 -> restarted worker, fixed"), exactly as Step 3 injected long-term memories.

TIP
Tip
Record failures as deliberately as successes. An agent that retrieves success=false episodes can avoid repeating a dead-end approach — often more valuable than knowing what worked, because the failure space is larger.
NOTE
Key Insight
For large logs, embed each episode into your vector store and retrieve by semantic similarity instead of keyword matching. Episodic and long-term memory then share one backend but answer different questions: episodic returns whole experiences, long-term returns distilled facts.

5 Semantic Memory: Facts and Knowledge

Semantic memory is the agent's store of facts true independent of any conversation or episode: business rules, product specs, policies, definitions. Unlike episodic memory there is no timestamp narrative — just structured truth the agent looks up.

For small, exact, frequently-changing facts, a structured table beats a vector store — you want deterministic lookups, not fuzzy similarity. Back it with a real data source, then expose it to the agent as a tool so the model fetches authoritative values instead of guessing. This is the Anthropic tool-use schema:

{
  "name": "lookup_policy",
  "description": "Return an authoritative company policy value. Always use this for refund windows, limits, and supported options instead of relying on memory.",
  "input_schema": {
    "type": "object",
    "properties": {
      "key": {
        "type": "string",
        "enum": ["refund_window_days", "max_file_upload_mb", "support_hours"]
      }
    },
    "required": ["key"]
  }
}

The handler is a trivial lookup — KNOWLEDGE.get(key, "UNKNOWN — escalate to a human") against a dict or database row. For larger, unstructured knowledge (docs, manuals, wikis), embed it into the same vector store you use for long-term memory and retrieve via RAG, using a metadata field so a policy lookup never returns a user preference.

WARNING
Watch Out
Never let the model improvise policy numbers from its training data — they will be plausible and wrong. Route every policy-dependent answer through a lookup_policy tool or RAG so the value is authoritative. The model's job is to reason, not to be the source of truth.
TIP
Tip
Semantic facts change. Keep the table in a database or config file you can update without redeploying, and version it. When a policy changes, you update one row — not a prompt buried in your codebase.

6 Combining All Four: A Support Agent

Now assemble a memory architecture for a support agent that remembers previous interactions and learns from resolved tickets. Each type plays a distinct role:

Need Memory type Mechanism
Follow the current chat Short-term Message list + compaction
Recall this customer's history and prefs Long-term Vector store, filtered by customer id
Reuse the fix from a similar past ticket Episodic Episode log, retrieved by similarity
Quote the correct refund policy Semantic lookup_policy tool

Each turn assembles context from all four before calling the model:

def support_turn(customer_id, user_msg, messages):
    messages = compact(messages)                                  # short-term
    messages.append({"role": "user", "content": user_msg})

    prefs = recall(f"customer {customer_id}: {user_msg}", k=3)     # long-term
    past = recent_successes(goal_keyword=user_msg.split()[0])      # episodic
    resolved = [f"{e['goal']} -> {e['outcome']}" for e in past]

    system = ("You are a support agent. Use lookup_policy for any policy value.\n"
              f"Known about this customer: {prefs}\n"
              f"Similar resolved tickets: {resolved}")

    resp = client.messages.create(
        model=MODEL, max_tokens=1024, system=system,
        tools=[POLICY_TOOL],          # the lookup_policy schema from Step 5
        messages=messages,
    )
    return resp, messages

When a ticket resolves, write back to memory so the agent improves — call record_episode(...) to log the whole run (episodic) and remember("customer X: resolved '...' via ...", metadata={...}) to distill the durable fact (long-term). This write-back loop is what makes it an agent rather than a stateless chatbot: each resolved ticket makes the next similar one easier.

NOTE
Key Insight
The four memory types compose, they don't compete. Short-term carries the live turn, long-term and episodic are retrieved into the system prompt, and semantic is fetched on demand via a tool. The write-back step closes the loop.
WARNING
Watch Out
Every retrieved memory you inject costs tokens and can mislead if irrelevant. Cap each retrieval (the k values), filter by metadata such as customer id, and measure whether memory actually improves outcomes — see Lesson 8 on evaluation before trusting it in production.

Questions & Answers

Q: Models keep shipping with bigger context windows. Won't that make external memory obsolete?
No. A larger window helps short-term memory but does nothing for persistence across sessions — close the process and the window is gone. It also doesn't fix cost (you pay per token resent) or the "lost in the middle" effect where models attend poorly to content buried in a huge prompt. Targeted retrieval stays cheaper and more reliable than dumping everything into the window.
Q: How do I stop the vector store from returning irrelevant or misleading memories?
Three levers. Curate writes — store distilled facts, not raw logs. Filter by metadata (customer id, type, recency) so a query searches only the relevant subset. And set a similarity threshold so you return nothing below it rather than always returning k results. An empty retrieval beats a confidently irrelevant one.
Q: When do I use a structured database versus a vector store for a memory?
Use a structured store (SQL, key-value) when lookups are exact — policy values, user settings, ticket status. Use a vector store when the query is fuzzy and you want semantic similarity — "find past tickets like this one." Most production agents use both: a database for facts and a vector store for experiences and documents.
Q: Doesn't summarizing short-term memory risk dropping something critical?
Yes — that is the tradeoff for staying within budget. Mitigate it by keeping the most recent turns verbatim, prompting the summarizer to preserve decisions and open questions explicitly, and writing anything truly durable to long-term memory before you compact. Treat the summary as a convenience, not the system of record.
Q: How do I keep memory from leaking one user's data into another user's session?
Partition every store by a tenant or user key and filter all retrievals on it. The customer id is in episode and memory metadata above; recall should pass that id as a metadata filter so a query can never match another customer's records. Redact secrets and PII at write time — once it's in the index, it's searchable.

Key Takeaways

  1. Memory is four subsystems, not one — short-term (the window), long-term (retrieved facts), episodic (recorded experiences), and semantic (authoritative truth). Design each separately.
  2. The context window is working memory — keep it dense with sliding windows, summarization, and token budgets; use prompt caching to make resends cheap.
  3. Long-term memory is RAG over your own history — embed, store, and retrieve top matches, but curate aggressively because a noisy store actively misleads the agent.
  4. Episodic memory teaches from experience — log episodes with outcomes, retrieve similar past runs, and record failures as deliberately as successes.
  5. Semantic facts go through tools, not the model's guesses — route policy and limit values through a lookup tool or RAG so answers stay authoritative.
  6. The write-back loop is what makes it an agent — distilling resolved work back into memory turns yesterday's experience into tomorrow's capability.

Next Steps: Lesson 5: Planning & Reasoning