Context engineering
Everything so far has been about a single call. A real assistant is a loop, and the loop's hard problem is not the prompt — it is deciding, turn after turn, which tokens deserve a place in the window. Context engineering is the discipline of curating the smallest set of high-signal tokens that produces the outcome, and it has four operations, one uncomfortable failure curve, and one live argument about whether a big enough window makes retrieval obsolete.
The smallest set of high-signal tokens
A budget with a goal, not a bucket with a limit
The context window is finite, attention is not free, and — this is the part that surprises people — accuracy degrades well inside the advertised limit. So the organising question for every turn is not "what could be relevant?" but "what is the smallest set of tokens that still produces this outcome?". That question has teeth because the two failure directions are asymmetric. Put too little in and the model guesses. Put too much in and it pays for every extra token in latency, in cost, and often in accuracy, because the signal is now diluted by near-miss documents and stale turns.
Anthropic's "Effective context engineering for AI agents" (29 September 2025) gives the taxonomy this part uses: write information down outside the window, select only what the current step needs, compress what must stay but need not stay verbatim, and isolate what should never share a context with anything privileged. Each is a concrete operation with a token cost you can measure, and each trades something away.
Write, select, compress, isolate
Four operations, four trades
Write moves state out of the window and into something durable — a scratchpad, a note table, a task list. The agent that writes down "the docs live under /ai/" does not need to keep the turn that discovered it. Select is retrieval brought to bear on the current step: pull five passages, not fifty. Compress summarises what must remain but need not remain verbatim, and it is the operation to be most suspicious of, because it is lossy and it is adversarial — an evicted fact is indistinguishable from a fact that never existed, so the agent silently redoes work or, worse, answers without it. Isolate keeps untrusted content away from privileged content and away from tools, which is the same move as the security part's quarantine and is just as much a context-budget decision as a safety one.
The diagram below is a prompt being assembled. Toggle each operation and watch which segment of the prompt it removes, what it adds back, and what the total costs.
Top: the four operations. Bottom: the resulting prompt as a single stacked bar — the running total is the cost of this context.
Context composition over many turns
Where the tokens actually go, step by step
On a single question the budget is easy. On a forty-step agent run it is the whole game: system instructions and tool definitions are a fixed tax on every step, retrieved passages recur, and the transcript is a term that grows without bound unless something removes it. The film strip below replays a scripted forty-step run through the shared docs-assistant scenario and shows the per-step composition of the prompt. Switch the compaction policy and watch both the total and a seeded answer-quality score move against each other.
Read the two policies carefully, because they fail differently. Eviction simply drops the oldest turns, which keeps the budget clean and loses facts — the run then re-retrieves them, and you can see the redo events in the readout. Structured notes keep an extracted fact table that survives compaction, so the run avoids most redos at a small per-fact token cost. That is the whole reason "write" sits first in the taxonomy: notes are how you compact without losing the things you cannot afford to lose.
One stacked column per step; the dashed line is the token ceiling. Redo and eviction events show up as changes in the shape, not as an error message.
Context rot and the lost middle
The window is not the same everywhere inside itself
Liu et al. (arXiv:2307.03172, TACL 2024) placed the answer to a question at varying depths of a long context and measured accuracy. The curve is a U: best when the relevant passage is at the very start (primacy) or the very end (recency), and worst in the middle. Chroma's Context Rot report (July 2025) found the same shape across many models, with degradation beginning well below the advertised window. A model does not read its context like a file; it reads it like a person reads a long document they were handed once.
This is why "just put everything in the window" is a strategy with a hidden accuracy cost. Each added passage spends budget and, if it lands mid-context, competes with the passage you actually needed. The playground below places a needle at a chosen depth and reads off a seeded recall rate interpolated through the measured bands — move the needle and watch the middle dip. The instruction-strength slider shows the caveat honestly: an explicit instruction pulls the whole curve up, but nothing in the current toolchain removes the dip entirely.
Top: a long context with the needle at the chosen depth. Bottom: seeded recall against needle position — the U-curve, with your position marked.
Long context versus retrieval
An argument with both sides on the table
If a model can attend to a million tokens, why build an index? The case for the long window is real: retrieval infrastructure is a second system to run, to keep fresh and to debug, and stuffing a small corpus into the prompt removes an entire class of "the retriever returned the wrong chunk" failures. For a handful of documents and a question whose evidence is spread thinly across all of them, the window genuinely is enough, and the pipeline is one call instead of four.
The case against is threefold. Cost and latency grow with the tokens you send, and attention cost grows faster than linearly. Freshness and permissions do not go away: a changed document and a per-user ACL still require a lookup before you decide what may be read. And accuracy is not monotonic in context length — the lost-middle curve from the previous section applies to the whole corpus you just stuffed in. So the honest position is a crossover, not a verdict: the window wins on small, stable, high-recall corpora; retrieval wins as the corpus grows, as queries repeat, and as access control matters. The panel below plots both against corpus size, in tokens per query on the left and answer quality on the right.
Left: tokens per query, log scale, with the advertised window as a ceiling. Right: seeded answer quality. The crossover is the point, not either endpoint.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| What is the goal of context engineering? | The smallest set of high-signal tokens that still produces the outcome — measured, not assumed. |
| What are the four operations? | Write (state outside the window), select (retrieve narrowly), compress (summarise), isolate (quarantine untrusted content). |
| Why is compaction dangerous? | It is lossy and adversarial: an evicted fact looks like a fact that never existed, so the agent silently redoes work. |
| Where in the window is evidence best read? | The start and the end. The middle is measurably worst — the lost-in-the-middle U-curve. |
| Does a big window make retrieval obsolete? | No. It removes retrieval for small, stable corpora; cost, freshness, ACLs and context rot still favour an index as the corpus grows. |
| What is just-in-time retrieval? | Keeping identifiers in context and fetching the body only when the current step needs it, rather than pre-loading everything. |
| What are structured notes for? | They survive compaction: facts written down outside the window can be re-read instead of re-derived. |
| What should be measured? | Tokens per step, drops and redos, answer quality against context size, and the position of the evidence you actually rely on. |
Further reading
- Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL 2024 — the U-shaped accuracy curve this part's playground is parameterised by.
- Chroma, "Context Rot: How Increasing Input Tokens Impacts LLM Performance", July 2025 — degradation across many models inside the advertised window.
- Anthropic, "Effective context engineering for AI agents", 29 September 2025 — the write/select/compress/isolate taxonomy.
- Anthropic, "Introducing Contextual Retrieval", 19 September 2024 — reported top-20 retrieval failure falling from 5.7% to 1.9% with contextual embeddings, BM25 and reranking.
- Li et al., "Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach", EMNLP 2024 industry track — long-context models win on average when resourced fully, but retrieval's far lower cost keeps it in the pipeline; their Self-Route routes each query to one or the other.
- Levy, Jacoby & Goldberg, "Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models", ACL 2024 — reasoning degrades at input lengths far below the technical maximum.
- Anthropic, "Building effective agents", December 2024 — when a workflow beats an open-ended loop, and why the simplest thing that works is usually right.