Caching-aware prompt layout
Every turn resends the whole prompt, and the provider charges for what it can recompute versus what it can reuse. Prefix caching is what lets it reuse, but the cache is keyed on the exact bytes at the front of your prompt. That makes layout — the order of the blocks you assemble — a cost and latency decision on the same footing as retrieval or model choice. This part is a puzzle: you reorder a real prompt, and 100 simulated requests tell you what it cost.
A hit is a longest-common-prefix match
The provider matches from the front, byte for byte
Prefill turns prompt tokens into key/value vectors, and it is the compute-heavy half of inference. When a provider offers prompt caching, it stores those vectors for a prefix it has already seen and lets the next request reuse them. The matching is not semantic: the provider hashes the prompt's leading blocks and looks for a byte-identical chain. The moment a block differs — a single character, a reordered JSON key, a new timestamp — the match stops, and every token after that point is recomputed at full price.
That is the whole lesson in one sentence: a cache hit is a longest-common-prefix match, so the reusable part of your prompt is whatever is byte-identical from the very first byte. Position matters more than content. A block that is stable but placed after a volatile block is worth nothing.
Anthropic's cache reads bill at 0.1× the base input rate and cache writes at 1.25×, with a 5-minute time-to-live by default; OpenAI's automatic cache reports it as cached_tokens at a 0.5× input rate and only engages once a prompt is over 1,024 tokens. The multipliers differ by vendor. What does not differ is the mechanism: the hit rate is a property of your layout, not of the vendor.
datetime.now() in the system block, a tool schema serialised with non-deterministic key order, a few-shot set shuffled into a different order each request, a stray trailing newline. None of them changes the answer. Every one of them resets the matched prefix to zero.Reorder the prompt, watch 100 requests
The load-bearing demo — click a block, then move it
Below is a seven-block prompt, started in the worst possible order: the timestamp sits at the very top. Move the blocks with the ▲/▼ arrows on each row (click the row first to select it), and toggle the two silent invalidators. Every change re-runs the same 100 simulated requests through AppSim.cacheSim with a 64-block prefix cache, and the strip along the bottom shows each request: the reused tokens in blue, the recomputed tokens in red, the pale remainder in grey.
Top: the prompt block stack, in send order. Bottom: 100 requests left to right; blue = cache read, red = recomputed after the first mismatch.
What one byte-identical change costs
A reordered key, a newline, a timestamp
The failure mode is not a bad prompt — it is a prompt that is almost identical to yesterday's. The bar below shows the matched prefix length for the same 7,460-token prompt under five scenarios. Every row keeps the same content; only the bytes at the boundary change. The matched prefix is everything the provider can reuse, and it collapses from the point of the first difference onward.
Solid = reused prefix, pale = recomputed at full price. The prompt is identical in every row; only the byte at the boundary changes.
The fix is unglamorous: serialise tool schemas with sorted keys, strip volatile fields out of cached blocks, pin a literal example order, and put anything time-dependent at the end. Cache breakpoints are explicit markers you can place to say "everything above this is stable"; the provider stores up to a few of them per request, so you can cache a shared system-plus-tools prefix and a per-user stable head without caching the volatile tail.
Write once, read many
Why the write premium is not a reason to skip caching
A cache write costs 1.25× the input rate and a cache read costs 0.1×, so the first request that populates the cache is slightly more expensive than no cache at all. The question is how many times the prefix is reused before the 5-minute TTL expires. If your system prefix is reused even twice, you are ahead — this is arithmetic, not judgement.
Drag the turns and the per-turn growth. The solid line is no cache; the dashed line is a warm cache that only pays read prices; the dotted line is the honest one, including the one-off write premium on the first turn.
Cumulative cost across turns, in USD. Solid = cold every turn, dashed = read-price only, dotted = with the 1.25× write premium included.
Measure cache_read_input_tokens, not vibes
If you cannot see the metric, you are paying for it
Every provider that caches exposes what actually happened. Anthropic returns cache_read_input_tokens and cache_creation_input_tokens; OpenAI returns cached_tokens inside the usage object. Those fields, not a dashboard's aggregate hit rate, are the ground truth for a single request — and the running sum of the read-price savings is the number you report to whoever pays the bill. The log below is twelve consecutive calls: the first four still had the timestamp on top, then the layout was fixed. The clicks are visible in the bill.
One column per call = cache_read_input_tokens. The dashed line marks the deploy that moved the timestamp to the tail; the light line is cumulative savings at the read price.
Cheat sheet
| Question | The answer |
|---|---|
| What does a prefix cache match? | The longest byte-identical prefix, from the first byte. A semantic near-match is a miss. |
| How should the prompt be ordered? | Stable content first (system, tools, examples), volatile content last (retrieval, history, timestamp, question). |
| What is a cache breakpoint? | An explicit marker that caches everything above it. You can have several per request. |
| What is the default TTL? | About 5 minutes at Anthropic; provider-specific. A gap longer than the TTL is a cold start. |
| What do read and write cost? | Read ≈ 0.1× base input; write ≈ 1.25×. Two reuses of a prefix already beat no caching. |
| What breaks the prefix? | A timestamp, a reordered JSON key, a shuffled example order, a stray newline — anything non-deterministic at the top. |
| What do I measure? | cache_read_input_tokens / cached_tokens per call, plus the running savings. Not a vibe. |
Further reading
- Anthropic, Prompt caching — the cache-write 1.25× and cache-read 0.1× multipliers, the 5-minute TTL, and the
cache_read_input_tokensusage fields this part instruments. - OpenAI, Prompt caching, October 2024 — automatic prefix matching over 1,024 tokens and the
cached_tokensusage field. - DeepSeek, Context caching — a disk-based prefix cache with published cache-hit and cache-miss prices, showing the multipliers are a vendor policy, not a law.
- Anthropic, "Effective context engineering for AI agents", 29 September 2025 — where a stable prefix sits in the write/select/compress/isolate taxonomy.
- Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention", SOSP 2023 — the block-level KV manager that makes provider prefix caching possible.