Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A stateless text function

In, out, and nothing in between

Strip away the framing and an LLM API call is a pure function: text in, text out. You send a list of messages; you get one completion back. The provider keeps no model-side state between your requests, so a "conversation" is not a conversation at all — it is a sequence of independent calls, each one re-sending everything the model is allowed to know.

That has an immediate consequence you can see in the request body. On turn three, the model is not recalling turns one and two; your application is pasting them back in. The transcript, the retrieved documents, the tool results, the system instructions — all of it is assembled fresh on every call, by your code, and priced as input tokens every time.

// what actually goes over the wire on turn 3 { "messages": [ { "role": "system", "content": "You are a docs assistant…" }, { "role": "user", "content": "What is the KV cache?" }, { "role": "assistant", "content": "It caches keys and values…" }, { "role": "user", "content": "And why does it grow?" } ] }
💡 The whole series in one sentence: the unit of engineering is not the prompt, it is the context window and the loop around it — the smallest set of high-signal tokens that produces the outcome, reassembled on every turn.

Step the turn counter below and watch the request body grow. The model's contribution never changes; the transcript is entirely yours.

One call per turn. Each call re-sends the entire transcript — the pale blocks are history the app pastes back in.

2

Where the “memory” actually lives

Three lanes, one of them yours

It helps to draw the system as three lanes. Your application is a stateful program: it owns a database, a transcript, a session store. The API in the middle is a stateless transport. The model is a stateless function. Only the leftmost lane accumulates anything between turns, and that is why every product decision about memory — what to keep, what to summarise, what to retrieve — is a decision your code makes, not the model's.

The picture below advances one turn at a time. The application lane gains a row each turn; the API and model lanes are redrawn from scratch every time.

Steady-state boxes on the right are transient; the growing column on the left is your state.

⚠️ Corollary: if your app “forgets”, the bug is in your state management, not the model. If your app “remembers something it should not”, the leak is in what your code chose to resend.
3

The context window is a budget

Not a bucket you fill until it breaks

The context window is the maximum number of tokens the model can attend to in one call — prompt plus completion. Everything the model knows about your request has to fit, and everything you put in competes for that space: instructions, examples, the transcript, retrieved documents, tool definitions, and the answer itself. The engineering move is to treat it as a budget with named line items, and to reserve room for the output before you spend the rest.

Drag the line items below. The bars are a single prompt's token composition against a fixed window; the readout says how much is left and warns before you overrun.

One stacked column per configuration. Hover the readout to see which line item is the largest.

4

Resending is quadratic, and caching is the brake

Why a long conversation gets expensive fast

Because every turn re-sends the whole transcript, the input tokens a provider sees are not linear in the number of turns — they are triangular. Ten turns of a growing conversation are not ten requests' worth of input; they are the sum of ten growing prefixes. Let each turn add h tokens and a conversation of N turns costs roughly h·N(N+1)/2 input tokens, before any retrieval or tool output. This is the arithmetic that makes agent loops expensive, and the reason prompt caching matters so much: if the prefix is byte-identical, a provider can reuse the cached computation and bill the repeat at a fraction of the input price.

The chart below shows cumulative billed input tokens over twenty turns, with and without prefix caching. Toggle caching and move the per-turn growth to see the gap widen.

Cumulative input tokens across 20 turns. Dashed = with a warm prefix cache; solid = cold every turn.

💡 Carry this forward: Parts 4 and 11 make the cost and cache story quantitative. Here it is enough to see that a loop over a growing transcript is the default, and that it is quadratic in turns unless something reuses the prefix.
5

Map of the two volumes

Where each remaining part fits

This volume takes the stateless function and builds upward: tokens and budgets (Part 2), sampling (3), cost and latency (4), and an eval harness you can run from the start (5). Act II is prompting and context design; Act III is retrieval. The companion volume, Agents in Action, continues from an agent that can search and read, through evaluation, fine-tuning and production. Where a topic belongs to the training or serving layer, both volumes link out rather than re-teach.

Cheat sheet

QuestionThe answer that shapes everything
Does the model remember my last call?No. It is a function of the text in this one call.
Where is the conversation stored?In your application. You assemble and re-send it every turn.
What is the context window?Prompt + completion, in tokens. A budget with line items, not a bucket.
What should I reserve first?Room for the output, before spending the rest on input.
Why does a long chat get expensive?Each turn re-sends the whole history: input cost is triangular in turns.
What slows that down?A byte-identical prefix across calls, so the provider can cache and bill it cheaper.
What is the unit of engineering?The context window and the loop around it — not the prompt.

Further reading

6

Check your understanding

0/4 answered