Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Quadratic prefill, linear KV

The two curves that decide everything

The training guide introduced prefill versus decode in a paragraph. Here the same distinction becomes a pair of curves you can price. A prefill pass over T tokens costs a linear weight term, 2NT, plus an attention term that grows with the square of the sequence:

$$\text{prefill FLOPs} \;\approx\; \underbrace{2NT}_{\text{weights}} \;+\; \underbrace{4\,L\,n_h\,d_h\,T^2}_{\text{QK}^\top + \text{AV}}$$

At 1,000 tokens the weight term dominates: you are paying to multiply the whole model over a short sequence. By roughly 30,000 tokens the terms cross, and past 100,000 the quadratic term is almost the entire bill. The KV cache is far better behaved — it is exactly linear, 2 L nkv dh T bytes — but it is also what decode must re-read on every token, so at long context decode stops being cheap in a way that is easy to miss until the bill arrives.

Top: prefill FLOPs, split into the linear and quadratic terms. Bottom: KV cache size, linear in context. Both on a log token axis.

💡 The honest read: on an H100-class part at a realistic 45% MFU, one 1M-token prefill of an 8B model is roughly 20 minutes of compute, while the KV cache produced alone is 131 GB — larger than the GPU. Long context is not a bigger window; it is a bigger data structure.
2

Context parallelism and ring attention

A window no single device can hold

When the KV cache exceeds one GPU, the sequence itself is sharded — that is context parallelism (CP), the long-context wing of the parallelism family covered in Part 12. Each device holds one contiguous block of queries and the keys/values for its own block. To attend to the whole sequence it must see every other block, so K/V blocks rotate around the ring: device i keeps its query block fixed while K/V blocks pass by, and after N hops every query has attended to every key. What makes this exact rather than approximate is that attention is computed in one streaming pass with an online softmax: each device maintains a running maximum m and denominator ℓ, rescaling the partial output whenever a visiting block raises the maximum.

Ring attention is the same online-softmax algebra FlashAttention uses inside one GPU, lifted across a network. Its appeal is that memory per device drops to O(T/N) and the only communication is one pass of K/V around the ring — overlappable with the block's own compute. Its cost is that CP does nothing for the quadratic FLOP term: eight devices split the cache, not the total attention work.

Press play. Each hop delivers one more K/V block to every device; watch the online-softmax maximum and denominator converge.

📌 Why this matters at 1M: the 131 GB KV cache above becomes 16 GB per device at N=8, which fits an H100 with room for weights. The ring buys capacity, not speed — and it is why CP is the one parallelism axis whose payoff grows with context length.
3

Chunked prefill at 1M tokens

One request, many iterations

A 1M-token prompt cannot prefill in a single scheduler iteration without stalling every other sequence for the duration. Chunked prefill splits it into fixed token budgets per step and interleaves those chunks with other requests' decode steps. The trade is explicit: a larger chunk finishes the prefill in fewer iterations (lower TTFT) but makes each iteration a long compute-bound matmul (higher inter-token latency for everyone else). The numbers below are the scheduler-facing version of the curves above — the same FLOPs, now spread across the step timeline.

Iterations to finish a 1M-token prefill at a fixed per-step token budget, and the ITL spike each iteration imposes.

4

Sparse attention at inference: Quest and DuoAttention

Do not compute the zeros

Dense causal attention still materialises a T × T score matrix. Three families of approximations avoid parts of it at inference time without retraining the model. Sliding window forgets everything older than W tokens. StreamingLLM adds a handful of sink tokens so the softmax always has a benign place to put mass. Quest is query-aware: it tracks a min/max key per page and, for each query, computes only the top-k critical pages. DuoAttention observes that only a small fraction of heads carry long-range information; retrieval heads keep the full KV, streaming heads keep sinks plus a constant window, and the two get different budgets.

A causal attention mask, rows = queries, columns = keys. Bright cells are computed; dim cells are skipped.

5

Eviction: sinks, H2O, SnapKV, PyramidKV

Keeping 12% of the cache and hoping it was the right 12%

Compression methods drop KV entries outright. The question that decides their usefulness is brutally simple: when the answer is buried in the middle of a 32K-token haystack, does the entry that holds it survive the budget? StreamingLLM keeps sinks plus a recent window. H2O keeps a heavy-hitter fraction by attention mass plus recent tokens. SnapKV votes on which prefix positions matter using an observation window at the end of the prompt. PyramidKV spends a layer-wise pyramidal budget, keeping more at the bottom layers where attention scatters. Slide the budget down and move the needle; the policies fail at different points.

One row per policy across 32K of KV. Lit cells are retained; the vertical marker is the needle.

6

Compression vs quantization vs eviction

Three levers, three failure modes

Quantization
2×
FP16 to FP8 stores each KV element in half the bytes. No entry is lost, so retrieval is preserved; the risk is numerical — long contexts accumulate more per-channel outliers, and Part 10's activation-outlier story is worse in the KV cache than in the weights.
Sparsity
2–7×
Skip attention to blocks likely to be irrelevant (Quest, sliding windows). Saves FLOPs, not memory — the cache still exists in full unless combined with eviction. Quality loss is query-dependent and hard to bound a priori.
Eviction
8×+
Delete entries permanently (H2O, SnapKV, PyramidKV). Saves memory and decode bandwidth, the actual bottleneck. The failure mode is silent: a dropped entry is unrecoverable for that request and often for its whole prefix cache.

They compose, and the composition is what production stacks ship: FP8 KV halves the bytes, a sliding window bounds most layers, DuoAttention leaves a few retrieval heads dense, and eviction trims what remains. Order matters. Quantizing first buys headroom for free and never loses a token, so it should be the default; eviction is the lever of last resort because it is the only one that can drop the answer. A useful mental model: quantization trades precision, sparsity trades FLOPs, and eviction trades recall — and recall is the one you cannot buy back at serving time.

⚠️ The prefix-cache trap. An evicted entry does not just hurt the current request — if the engine caches the compressed prefix, the next request that reuses it inherits the loss. Eviction policies interact with Part 8's radix cache in ways that are easy to miss in a single-request benchmark.
7

Measuring the damage

Needle-in-a-haystack is not enough

The standard long-context eval — plant a sentence at depth d in a context of length L, ask for it back — is a single-retrieval probe. It is necessary and nowhere near sufficient: models that score 100 on it still fail multi-hop questions, aggregation over many documents, and tasks whose relevant evidence is distributed rather than pinpointed. RULER and similar suites add multi-needle, variable-tracking and aggregation tracks precisely because eviction and sparsity often pass the single needle and fail everything else.

The heatmap below is a deterministic reconstruction of the shape these results take: accuracy degrades with context length, drops fastest in the middle depths where neither recency nor the sink protects an entry, and is the first thing a KV budget kills. The practical rule is to run the eval at your serving budget, not at full cache, and to report the depth profile rather than one average.

Retrieval accuracy by needle depth and context length. Brighter is better.

8

The honest cost per request

A 200K-token agent turn

Put it all together and price the workload that actually exercises long context: an agent turn that reads 200K tokens of retrieved context and emits 2,000 tokens of plan, then does it again on the next turn. Prefill is a one-time quadratic charge; decode is a per-token linear re-read of a 26 GB cache; and if the prefix is reused across turns, the second turn's prefill collapses to near zero while its decode cost does not move at all. That asymmetry — reusable prefill, irreducible decode — is the economics of long-context agents, and it is why extending the trained window and serving that window are different problems: here the window exists, and the question is what it costs to serve.

Where one agent turn's GPU-seconds go, and what they cost at a nominal rate of USD 2.50 per GPU-hour.

Cheat sheet

LeverWhat it costsWhat it savesFailure mode
Bigger window (training)Quadratic prefill, linear KV—Adds serving bill, not capability
Context parallelism / ringK/V pass around the ring, CP collectivesMemory per device O(T/N)FLOPs unchanged; needs fast interconnect
Chunked prefillMore iterations, ITL spikesBounded TTFT, other requests keep movingSmall chunks starve throughput
Sliding windowForgets older keysDecode O(W), bounded cacheLong-range retrieval lost past W
Quest top-k pagesQuery-aware page scoringUp to 7.03× latency (2.23× attention)Misses evidence the score underestimates
DuoAttentionHead classification at load timeUp to 2.18× decode, 1.73× prefillA misclassified retrieval head loses context
H2O / SnapKV / PyramidKVIrreversible eviction8×+ memory, direct decode bandwidthDrops the needle; poisons cached prefixes
FP8 KV~0–0.5% quality2× cache bytes, 2× decode headroomOutlier sensitivity at long context

Further reading

?

Check your understanding

0/5 answered