Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A hit is a longest-common-prefix match

The provider matches from the front, byte for byte

Prefill turns prompt tokens into key/value vectors, and it is the compute-heavy half of inference. When a provider offers prompt caching, it stores those vectors for a prefix it has already seen and lets the next request reuse them. The matching is not semantic: the provider hashes the prompt's leading blocks and looks for a byte-identical chain. The moment a block differs — a single character, a reordered JSON key, a new timestamp — the match stops, and every token after that point is recomputed at full price.

That is the whole lesson in one sentence: a cache hit is a longest-common-prefix match, so the reusable part of your prompt is whatever is byte-identical from the very first byte. Position matters more than content. A block that is stable but placed after a volatile block is worth nothing.

// the layout that hits system prompt // stable, identical every call tool definitions // stable, identical every call few-shot examples // stable, identical every call ──────────────────── ← cache breakpoint: everything above is reused retrieved documents // volatile: changes with the query conversation history // volatile: grows every turn timestamp // volatile: datetime.now() user question // volatile: different every call
💡 The rule: stable content first, volatile content last. The reusable prefix is worth the most when it is long, and it is cheap to make long — system instructions, tool schemas and few-shot examples are exactly the blocks you control.

Anthropic's cache reads bill at 0.1× the base input rate and cache writes at 1.25×, with a 5-minute time-to-live by default; OpenAI's automatic cache reports it as cached_tokens at a 0.5× input rate and only engages once a prompt is over 1,024 tokens. The multipliers differ by vendor. What does not differ is the mechanism: the hit rate is a property of your layout, not of the vendor.

⚠️ The silent invalidators: a datetime.now() in the system block, a tool schema serialised with non-deterministic key order, a few-shot set shuffled into a different order each request, a stray trailing newline. None of them changes the answer. Every one of them resets the matched prefix to zero.
2

Reorder the prompt, watch 100 requests

The load-bearing demo — click a block, then move it

Below is a seven-block prompt, started in the worst possible order: the timestamp sits at the very top. Move the blocks with the ▲/▼ arrows on each row (click the row first to select it), and toggle the two silent invalidators. Every change re-runs the same 100 simulated requests through AppSim.cacheSim with a 64-block prefix cache, and the strip along the bottom shows each request: the reused tokens in blue, the recomputed tokens in red, the pale remainder in grey.

Top: the prompt block stack, in send order. Bottom: 100 requests left to right; blue = cache read, red = recomputed after the first mismatch.

💡 What to watch: with the timestamp on top the hit rate is 0%, because the first byte differs on every call. Move the three stable blocks to the front and it jumps to 99% — the same content, rearranged, with a large real cost difference.
3

What one byte-identical change costs

A reordered key, a newline, a timestamp

The failure mode is not a bad prompt — it is a prompt that is almost identical to yesterday's. The bar below shows the matched prefix length for the same 7,460-token prompt under five scenarios. Every row keeps the same content; only the bytes at the boundary change. The matched prefix is everything the provider can reuse, and it collapses from the point of the first difference onward.

Solid = reused prefix, pale = recomputed at full price. The prompt is identical in every row; only the byte at the boundary changes.

The fix is unglamorous: serialise tool schemas with sorted keys, strip volatile fields out of cached blocks, pin a literal example order, and put anything time-dependent at the end. Cache breakpoints are explicit markers you can place to say "everything above this is stable"; the provider stores up to a few of them per request, so you can cache a shared system-plus-tools prefix and a per-user stable head without caching the volatile tail.

4

Write once, read many

Why the write premium is not a reason to skip caching

A cache write costs 1.25× the input rate and a cache read costs 0.1×, so the first request that populates the cache is slightly more expensive than no cache at all. The question is how many times the prefix is reused before the 5-minute TTL expires. If your system prefix is reused even twice, you are ahead — this is arithmetic, not judgement.

Drag the turns and the per-turn growth. The solid line is no cache; the dashed line is a warm cache that only pays read prices; the dotted line is the honest one, including the one-off write premium on the first turn.

Cumulative cost across turns, in USD. Solid = cold every turn, dashed = read-price only, dotted = with the 1.25× write premium included.

5

Measure cache_read_input_tokens, not vibes

If you cannot see the metric, you are paying for it

Every provider that caches exposes what actually happened. Anthropic returns cache_read_input_tokens and cache_creation_input_tokens; OpenAI returns cached_tokens inside the usage object. Those fields, not a dashboard's aggregate hit rate, are the ground truth for a single request — and the running sum of the read-price savings is the number you report to whoever pays the bill. The log below is twelve consecutive calls: the first four still had the timestamp on top, then the layout was fixed. The clicks are visible in the bill.

One column per call = cache_read_input_tokens. The dashed line marks the deploy that moved the timestamp to the tail; the light line is cumulative savings at the read price.

💡 The instrument rule: log the per-call cache fields next to the prompt layout hash. A hit rate that drops is almost always a layout change or a TTL gap — not the model, and not the users.

Cheat sheet

QuestionThe answer
What does a prefix cache match?The longest byte-identical prefix, from the first byte. A semantic near-match is a miss.
How should the prompt be ordered?Stable content first (system, tools, examples), volatile content last (retrieval, history, timestamp, question).
What is a cache breakpoint?An explicit marker that caches everything above it. You can have several per request.
What is the default TTL?About 5 minutes at Anthropic; provider-specific. A gap longer than the TTL is a cold start.
What do read and write cost?Read ≈ 0.1× base input; write ≈ 1.25×. Two reuses of a prefix already beat no caching.
What breaks the prefix?A timestamp, a reordered JSON key, a shuffled example order, a stray newline — anything non-deterministic at the top.
What do I measure?cache_read_input_tokens / cached_tokens per call, plus the running savings. Not a vibe.

Further reading

6

Check your understanding

0/5 answered