What you're actually building on
Before any prompt, any retrieval, any agent, there is one fact that reorganises everything else: the model is a stateless function of the text you hand it. It does not remember your last request, it has no session, and it cannot see your database. Everything that looks like memory is your own code deciding what to resend. This part fixes that picture, and the context window that follows from it, before the rest of the two volumes build on top.
A stateless text function
In, out, and nothing in between
Strip away the framing and an LLM API call is a pure function: text in, text out. You send a list of messages; you get one completion back. The provider keeps no model-side state between your requests, so a "conversation" is not a conversation at all — it is a sequence of independent calls, each one re-sending everything the model is allowed to know.
That has an immediate consequence you can see in the request body. On turn three, the model is not recalling turns one and two; your application is pasting them back in. The transcript, the retrieved documents, the tool results, the system instructions — all of it is assembled fresh on every call, by your code, and priced as input tokens every time.
Step the turn counter below and watch the request body grow. The model's contribution never changes; the transcript is entirely yours.
One call per turn. Each call re-sends the entire transcript — the pale blocks are history the app pastes back in.
Where the “memory” actually lives
Three lanes, one of them yours
It helps to draw the system as three lanes. Your application is a stateful program: it owns a database, a transcript, a session store. The API in the middle is a stateless transport. The model is a stateless function. Only the leftmost lane accumulates anything between turns, and that is why every product decision about memory — what to keep, what to summarise, what to retrieve — is a decision your code makes, not the model's.
The picture below advances one turn at a time. The application lane gains a row each turn; the API and model lanes are redrawn from scratch every time.
Steady-state boxes on the right are transient; the growing column on the left is your state.
The context window is a budget
Not a bucket you fill until it breaks
The context window is the maximum number of tokens the model can attend to in one call — prompt plus completion. Everything the model knows about your request has to fit, and everything you put in competes for that space: instructions, examples, the transcript, retrieved documents, tool definitions, and the answer itself. The engineering move is to treat it as a budget with named line items, and to reserve room for the output before you spend the rest.
Drag the line items below. The bars are a single prompt's token composition against a fixed window; the readout says how much is left and warns before you overrun.
One stacked column per configuration. Hover the readout to see which line item is the largest.
Resending is quadratic, and caching is the brake
Why a long conversation gets expensive fast
Because every turn re-sends the whole transcript, the input tokens a provider sees are not linear in the number of turns — they are triangular. Ten turns of a growing conversation are not ten requests' worth of input; they are the sum of ten growing prefixes. Let each turn add h tokens and a conversation of N turns costs roughly h·N(N+1)/2 input tokens, before any retrieval or tool output. This is the arithmetic that makes agent loops expensive, and the reason prompt caching matters so much: if the prefix is byte-identical, a provider can reuse the cached computation and bill the repeat at a fraction of the input price.
The chart below shows cumulative billed input tokens over twenty turns, with and without prefix caching. Toggle caching and move the per-turn growth to see the gap widen.
Cumulative input tokens across 20 turns. Dashed = with a warm prefix cache; solid = cold every turn.
Map of the two volumes
Where each remaining part fits
This volume takes the stateless function and builds upward: tokens and budgets (Part 2), sampling (3), cost and latency (4), and an eval harness you can run from the start (5). Act II is prompting and context design; Act III is retrieval. The companion volume, Agents in Action, continues from an agent that can search and read, through evaluation, fine-tuning and production. Where a topic belongs to the training or serving layer, both volumes link out rather than re-teach.
Cheat sheet
| Question | The answer that shapes everything |
|---|---|
| Does the model remember my last call? | No. It is a function of the text in this one call. |
| Where is the conversation stored? | In your application. You assemble and re-send it every turn. |
| What is the context window? | Prompt + completion, in tokens. A budget with line items, not a bucket. |
| What should I reserve first? | Room for the output, before spending the rest on input. |
| Why does a long chat get expensive? | Each turn re-sends the whole history: input cost is triangular in turns. |
| What slows that down? | A byte-identical prefix across calls, so the provider can cache and bill it cheaper. |
| What is the unit of engineering? | The context window and the loop around it — not the prompt. |
Further reading
- Anthropic, "Building effective agents", December 2024 — the workflow-versus-agent split and why the simplest thing that works is usually right.
- Anthropic, "Effective context engineering for AI agents", 29 September 2025 — the write/select/compress/isolate taxonomy this series restates as its first thesis.
- Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL 2024 — the U-shaped accuracy curve behind Part 10's argument about context rot.
- OpenAI, Prompt caching guide, October 2024 — how prefix reuse is billed and why prompt layout is a cost decision.
- Anthropic, Prompt caching — cache-write and cache-read multipliers, and the 5-minute TTL that shapes Part 11.