Memory and compaction
An agent that runs for forty steps has a memory problem disguised as a context problem. The window fills with history, so something has to leave — and whatever leaves takes its knowledge with it. This part separates the four things people mean by "memory", walks through the three mechanisms that manage it, and then watches a fact get evicted from a live run and quietly re-derived three steps later. It ends on the uncomfortable property that makes compaction dangerous rather than merely lossy.
Four things called memory
Working, episodic, semantic, procedural
"Memory" in agent design is four different storage problems with four different answers, and most confusion comes from solving one while asking for another. Working memory is the live context window: everything resident and attended to right now, and by construction the smallest and most expensive store. Episodic memory is the record of what happened — the transcript of past sessions, retrievable later, useful for "what did we try last time?". Semantic memory is durable fact: the user's name, their timezone, the fact that the corpus lives under /ai/. Procedural memory is knowing how — a system prompt, a skill body, a fine-tuned reflex — which is behaviour rather than data.
The distinction is practical because the answers diverge. Names and preferences belong in a fact store that outlives the session and is looked up when relevant. A task in progress belongs in the working context and nowhere else. A procedure belongs in a prompt or a skill, where it costs tokens once rather than being re-derived. Pick a need below and see which store it actually belongs to.
The need on the left; the store that should hold it on the right. Getting this wrong is how an agent "remembers" a stale preference forever.
Compaction, context editing, external notes
Three mechanisms with three different failure modes
Compaction is summarisation: replace a span of history with a shorter paraphrase so the gist survives. It is the default in most agent harnesses because it is easy, and it is the one to be most suspicious of, because the model writing the summary decides what mattered. Context editing is surgery on specific spans: drop the oldest tool results, clear the raw documents after their content has been extracted, remove a stale branch of the loop. It is cheaper and more predictable than summarisation and loses more, deliberately. External notes invert the problem: the agent writes durable structured state outside the window — a fact table, a scratch file, a task list — and re-reads the part it needs. Nothing is compacted, because the aggregate was never in context to begin with.
The research literature is largely about that third mechanism. MemGPT (Packer et al., arXiv:2310.08560) framed the context window as fast memory and the rest as slow storage, with the model itself paging material in and out through function calls — the operating-system metaphor that Letta now productises. Mem0 (arXiv:2504.19413) builds a lighter pipeline: extract salient facts from a conversation, consolidate them against what is already stored, and retrieve the relevant subset later. Zep/Graphiti (arXiv:2501.13956) adds time: a temporal knowledge graph in which a fact carries a validity interval, so a superseded fact is invalidated rather than deleted and a query can be scoped to "as of last week".
The mechanisms differ in what they trust. Compaction trusts a summary. Context editing trusts a rule. Notes trust structure. When exact values matter — an order id, a price, a file path — only the last of those has an answer that is guaranteed to be the original.
The context film strip
Watch a fact get dropped, then watch it get re-derived
Here is a scripted forty-step docs-assistant run with eight facts it learns along the way. Each column is one step's prompt, stacked by segment; the dashed line is the ceiling. Switch the compaction policy and scrub through the run. The triangles mark steps where a fact was evicted; the diamonds mark steps where the agent had to redo work because something it needed was gone.
The scenario is deliberate about which facts are fragile. f3, "the corpus has sixty documents", has low salience — so summarisation drops it, and eviction drops it sooner — and it is needed again later. In the default eviction run it is evicted at step 18 and re-derived at step 21. Three steps of work, paid twice, with nothing in the transcript to say so.
One stacked column per step. ▼ eviction · ◆ redo. Scrub to step through and read the fact labels in the panel.
The compaction tax, measured
Cheapest context is not the best answer
The four policies in the same run, at the same 12,000-token ceiling, priced three ways: total tokens sent across the whole run, the number of redos the policy caused, and a seeded task-success score. The tradeoff is not a single axis. "No compaction" wins on redos — nothing is ever evicted, so nothing is ever re-derived — and loses badly anyway, because on 28 of the 40 steps the prompt runs past the ceiling, which in a real deployment means silent truncation. Eviction sends the fewest tokens and causes the most redos. Summarisation lands in the middle. Structured notes send a little more than eviction and cause none.
Read the bars as a reminder that "how many tokens did it cost" is not the objective function. The objective is the task, and the token metric is only a good proxy when it is not buying redos.
Three panels, four policies each. The best value in each panel is marked; no policy is best in all three.
Notes versus summarisation
Smaller context, zero redos — for a maintenance price
The cleanest comparison is between the two policies that both keep the run inside the ceiling: summarisation, which compresses history into prose, and structured notes, which moves facts out of history altogether. At a tight ceiling the notes run sends less and redoes nothing. Slide the ceiling up and the token totals converge — the advantage is not that notes are magically cheaper, it is that notes never make the agent re-derive something it already knew.
The price is maintenance, and it is real. Notes need a schema, a place to live, a policy for updating them, and a rule for when they are stale. A note that says the corpus has sixty documents is a liability the moment someone adds a sixty-first. Summarisation has no such bookkeeping; it just cannot be trusted with the exact value.
Prompt tokens per step, same run, two policies. Ticks on the summary line are redos caused by a fact that had been compressed away.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| What is working memory? | The live context window — resident, attended to, and the most expensive store you have. |
| What is episodic memory? | The record of what happened: past sessions, retrievable later for "what did we try last time". |
| What is semantic memory? | Durable facts about the world and the user, stored outside the window and looked up when relevant. |
| What is procedural memory? | How-to: a system prompt, a skill body, a tuned reflex. Behaviour, not data. |
| Compaction versus context editing? | Compaction paraphrases a span; context editing deletes specific spans by rule. Editing is predictable and loses more. |
| What do external notes buy? | Facts move out of history, so they are re-read by key instead of re-derived — and never depend on a summary's judgement. |
| Why is compaction adversarial? | An evicted fact is indistinguishable from a fact that never existed, so the agent silently redoes work. |
| What does MemGPT page? | Material between a fast main context and slow external storage, with the model driving the paging through function calls. |
| What does a temporal graph add? | A validity interval per fact, so a superseded fact is invalidated rather than deleted and queries can be time-scoped. |
| What should be measured? | Tokens per step, drops, redos and task success — not tokens alone. |
Further reading
- Packer et al., "MemGPT: Towards LLMs as Operating Systems", 2023 — the paging metaphor, with the model managing its own memory through function calls.
- Chhikara et al., "Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory", 2025 — extraction, consolidation and retrieval of salient facts, with a graph variant.
- Rasmussen et al., "Zep: A Temporal Knowledge Graph Architecture for Agent Memory", 2025 — validity intervals so facts are superseded rather than deleted.
- Anthropic, "Managing context on the Claude Developer Platform", 2025 — context editing and a file-based memory tool: the external-notes mechanism as a platform feature.
- Anthropic, "Effective context engineering for AI agents", 29 September 2025 — write/select/compress/isolate, and why "write" is listed first.
- Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL 2024 — why a larger window does not remove the need to curate it.