Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The smallest set of high-signal tokens

A budget with a goal, not a bucket with a limit

The context window is finite, attention is not free, and — this is the part that surprises people — accuracy degrades well inside the advertised limit. So the organising question for every turn is not "what could be relevant?" but "what is the smallest set of tokens that still produces this outcome?". That question has teeth because the two failure directions are asymmetric. Put too little in and the model guesses. Put too much in and it pays for every extra token in latency, in cost, and often in accuracy, because the signal is now diluted by near-miss documents and stale turns.

Anthropic's "Effective context engineering for AI agents" (29 September 2025) gives the taxonomy this part uses: write information down outside the window, select only what the current step needs, compress what must stay but need not stay verbatim, and isolate what should never share a context with anything privileged. Each is a concrete operation with a token cost you can measure, and each trades something away.

💡 The thesis, in one sentence: context is not storage — it is a working set. The engineering move is to decide what is resident right now, and to treat everything else as something you can retrieve, summarise or re-derive later.
2

Write, select, compress, isolate

Four operations, four trades

Write moves state out of the window and into something durable — a scratchpad, a note table, a task list. The agent that writes down "the docs live under /ai/" does not need to keep the turn that discovered it. Select is retrieval brought to bear on the current step: pull five passages, not fifty. Compress summarises what must remain but need not remain verbatim, and it is the operation to be most suspicious of, because it is lossy and it is adversarial — an evicted fact is indistinguishable from a fact that never existed, so the agent silently redoes work or, worse, answers without it. Isolate keeps untrusted content away from privileged content and away from tools, which is the same move as the security part's quarantine and is just as much a context-budget decision as a safety one.

The diagram below is a prompt being assembled. Toggle each operation and watch which segment of the prompt it removes, what it adds back, and what the total costs.

Top: the four operations. Bottom: the resulting prompt as a single stacked bar — the running total is the cost of this context.

3

Context composition over many turns

Where the tokens actually go, step by step

On a single question the budget is easy. On a forty-step agent run it is the whole game: system instructions and tool definitions are a fixed tax on every step, retrieved passages recur, and the transcript is a term that grows without bound unless something removes it. The film strip below replays a scripted forty-step run through the shared docs-assistant scenario and shows the per-step composition of the prompt. Switch the compaction policy and watch both the total and a seeded answer-quality score move against each other.

Read the two policies carefully, because they fail differently. Eviction simply drops the oldest turns, which keeps the budget clean and loses facts — the run then re-retrieves them, and you can see the redo events in the readout. Structured notes keep an extracted fact table that survives compaction, so the run avoids most redos at a small per-fact token cost. That is the whole reason "write" sits first in the taxonomy: notes are how you compact without losing the things you cannot afford to lose.

One stacked column per step; the dashed line is the token ceiling. Redo and eviction events show up as changes in the shape, not as an error message.

⚠️ The compaction tax: compaction is lossy and adversarial. An evicted fact is indistinguishable from a fact that never existed, so the agent silently redoes work — the redo count in the readout is the price of the eviction, and it is invisible unless you measure it.
4

Context rot and the lost middle

The window is not the same everywhere inside itself

Liu et al. (arXiv:2307.03172, TACL 2024) placed the answer to a question at varying depths of a long context and measured accuracy. The curve is a U: best when the relevant passage is at the very start (primacy) or the very end (recency), and worst in the middle. Chroma's Context Rot report (July 2025) found the same shape across many models, with degradation beginning well below the advertised window. A model does not read its context like a file; it reads it like a person reads a long document they were handed once.

This is why "just put everything in the window" is a strategy with a hidden accuracy cost. Each added passage spends budget and, if it lands mid-context, competes with the passage you actually needed. The playground below places a needle at a chosen depth and reads off a seeded recall rate interpolated through the measured bands — move the needle and watch the middle dip. The instruction-strength slider shows the caveat honestly: an explicit instruction pulls the whole curve up, but nothing in the current toolchain removes the dip entirely.

Top: a long context with the needle at the chosen depth. Bottom: seeded recall against needle position — the U-curve, with your position marked.

5

Long context versus retrieval

An argument with both sides on the table

If a model can attend to a million tokens, why build an index? The case for the long window is real: retrieval infrastructure is a second system to run, to keep fresh and to debug, and stuffing a small corpus into the prompt removes an entire class of "the retriever returned the wrong chunk" failures. For a handful of documents and a question whose evidence is spread thinly across all of them, the window genuinely is enough, and the pipeline is one call instead of four.

The case against is threefold. Cost and latency grow with the tokens you send, and attention cost grows faster than linearly. Freshness and permissions do not go away: a changed document and a per-user ACL still require a lookup before you decide what may be read. And accuracy is not monotonic in context length — the lost-middle curve from the previous section applies to the whole corpus you just stuffed in. So the honest position is a crossover, not a verdict: the window wins on small, stable, high-recall corpora; retrieval wins as the corpus grows, as queries repeat, and as access control matters. The panel below plots both against corpus size, in tokens per query on the left and answer quality on the right.

Left: tokens per query, log scale, with the advertised window as a ceiling. Right: seeded answer quality. The crossover is the point, not either endpoint.

Cheat sheet

QuestionThe answer that shapes the build
What is the goal of context engineering?The smallest set of high-signal tokens that still produces the outcome — measured, not assumed.
What are the four operations?Write (state outside the window), select (retrieve narrowly), compress (summarise), isolate (quarantine untrusted content).
Why is compaction dangerous?It is lossy and adversarial: an evicted fact looks like a fact that never existed, so the agent silently redoes work.
Where in the window is evidence best read?The start and the end. The middle is measurably worst — the lost-in-the-middle U-curve.
Does a big window make retrieval obsolete?No. It removes retrieval for small, stable corpora; cost, freshness, ACLs and context rot still favour an index as the corpus grows.
What is just-in-time retrieval?Keeping identifiers in context and fetching the body only when the current step needs it, rather than pre-loading everything.
What are structured notes for?They survive compaction: facts written down outside the window can be re-read instead of re-derived.
What should be measured?Tokens per step, drops and redos, answer quality against context size, and the position of the evidence you actually rely on.

Further reading

6

Check your understanding

0/5 answered