Long-context serving
Extending the window is a training problem. Serving it is an arithmetic one: prefill compute grows with the square of the context while the KV cache grows linearly, and every long-context system in production is a negotiated bargain between the two. This part prices that bargain, from a single 1M-token request to the sparse-attention and eviction tricks that make it affordable.
Quadratic prefill, linear KV
The two curves that decide everything
The training guide introduced prefill versus decode in a paragraph. Here the same distinction becomes a pair of curves you can price. A prefill pass over T tokens costs a linear weight term, 2NT, plus an attention term that grows with the square of the sequence:
At 1,000 tokens the weight term dominates: you are paying to multiply the whole model over a short sequence. By roughly 30,000 tokens the terms cross, and past 100,000 the quadratic term is almost the entire bill. The KV cache is far better behaved — it is exactly linear, 2 L nkv dh T bytes — but it is also what decode must re-read on every token, so at long context decode stops being cheap in a way that is easy to miss until the bill arrives.
Top: prefill FLOPs, split into the linear and quadratic terms. Bottom: KV cache size, linear in context. Both on a log token axis.
Context parallelism and ring attention
A window no single device can hold
When the KV cache exceeds one GPU, the sequence itself is sharded — that is context parallelism (CP), the long-context wing of the parallelism family covered in Part 12. Each device holds one contiguous block of queries and the keys/values for its own block. To attend to the whole sequence it must see every other block, so K/V blocks rotate around the ring: device i keeps its query block fixed while K/V blocks pass by, and after N hops every query has attended to every key. What makes this exact rather than approximate is that attention is computed in one streaming pass with an online softmax: each device maintains a running maximum m and denominator ℓ, rescaling the partial output whenever a visiting block raises the maximum.
Ring attention is the same online-softmax algebra FlashAttention uses inside one GPU, lifted across a network. Its appeal is that memory per device drops to O(T/N) and the only communication is one pass of K/V around the ring — overlappable with the block's own compute. Its cost is that CP does nothing for the quadratic FLOP term: eight devices split the cache, not the total attention work.
Press play. Each hop delivers one more K/V block to every device; watch the online-softmax maximum and denominator converge.
Chunked prefill at 1M tokens
One request, many iterations
A 1M-token prompt cannot prefill in a single scheduler iteration without stalling every other sequence for the duration. Chunked prefill splits it into fixed token budgets per step and interleaves those chunks with other requests' decode steps. The trade is explicit: a larger chunk finishes the prefill in fewer iterations (lower TTFT) but makes each iteration a long compute-bound matmul (higher inter-token latency for everyone else). The numbers below are the scheduler-facing version of the curves above — the same FLOPs, now spread across the step timeline.
Iterations to finish a 1M-token prefill at a fixed per-step token budget, and the ITL spike each iteration imposes.
Sparse attention at inference: Quest and DuoAttention
Do not compute the zeros
Dense causal attention still materialises a T × T score matrix. Three families of approximations avoid parts of it at inference time without retraining the model. Sliding window forgets everything older than W tokens. StreamingLLM adds a handful of sink tokens so the softmax always has a benign place to put mass. Quest is query-aware: it tracks a min/max key per page and, for each query, computes only the top-k critical pages. DuoAttention observes that only a small fraction of heads carry long-range information; retrieval heads keep the full KV, streaming heads keep sinks plus a constant window, and the two get different budgets.
A causal attention mask, rows = queries, columns = keys. Bright cells are computed; dim cells are skipped.
Eviction: sinks, H2O, SnapKV, PyramidKV
Keeping 12% of the cache and hoping it was the right 12%
Compression methods drop KV entries outright. The question that decides their usefulness is brutally simple: when the answer is buried in the middle of a 32K-token haystack, does the entry that holds it survive the budget? StreamingLLM keeps sinks plus a recent window. H2O keeps a heavy-hitter fraction by attention mass plus recent tokens. SnapKV votes on which prefix positions matter using an observation window at the end of the prompt. PyramidKV spends a layer-wise pyramidal budget, keeping more at the bottom layers where attention scatters. Slide the budget down and move the needle; the policies fail at different points.
One row per policy across 32K of KV. Lit cells are retained; the vertical marker is the needle.
Compression vs quantization vs eviction
Three levers, three failure modes
They compose, and the composition is what production stacks ship: FP8 KV halves the bytes, a sliding window bounds most layers, DuoAttention leaves a few retrieval heads dense, and eviction trims what remains. Order matters. Quantizing first buys headroom for free and never loses a token, so it should be the default; eviction is the lever of last resort because it is the only one that can drop the answer. A useful mental model: quantization trades precision, sparsity trades FLOPs, and eviction trades recall — and recall is the one you cannot buy back at serving time.
Measuring the damage
Needle-in-a-haystack is not enough
The standard long-context eval — plant a sentence at depth d in a context of length L, ask for it back — is a single-retrieval probe. It is necessary and nowhere near sufficient: models that score 100 on it still fail multi-hop questions, aggregation over many documents, and tasks whose relevant evidence is distributed rather than pinpointed. RULER and similar suites add multi-needle, variable-tracking and aggregation tracks precisely because eviction and sparsity often pass the single needle and fail everything else.
The heatmap below is a deterministic reconstruction of the shape these results take: accuracy degrades with context length, drops fastest in the middle depths where neither recency nor the sink protects an entry, and is the first thing a KV budget kills. The practical rule is to run the eval at your serving budget, not at full cache, and to report the depth profile rather than one average.
Retrieval accuracy by needle depth and context length. Brighter is better.
The honest cost per request
A 200K-token agent turn
Put it all together and price the workload that actually exercises long context: an agent turn that reads 200K tokens of retrieved context and emits 2,000 tokens of plan, then does it again on the next turn. Prefill is a one-time quadratic charge; decode is a per-token linear re-read of a 26 GB cache; and if the prefix is reused across turns, the second turn's prefill collapses to near zero while its decode cost does not move at all. That asymmetry — reusable prefill, irreducible decode — is the economics of long-context agents, and it is why extending the trained window and serving that window are different problems: here the window exists, and the question is what it costs to serve.
Where one agent turn's GPU-seconds go, and what they cost at a nominal rate of USD 2.50 per GPU-hour.
Cheat sheet
| Lever | What it costs | What it saves | Failure mode |
|---|---|---|---|
| Bigger window (training) | Quadratic prefill, linear KV | — | Adds serving bill, not capability |
| Context parallelism / ring | K/V pass around the ring, CP collectives | Memory per device O(T/N) | FLOPs unchanged; needs fast interconnect |
| Chunked prefill | More iterations, ITL spikes | Bounded TTFT, other requests keep moving | Small chunks starve throughput |
| Sliding window | Forgets older keys | Decode O(W), bounded cache | Long-range retrieval lost past W |
| Quest top-k pages | Query-aware page scoring | Up to 7.03× latency (2.23× attention) | Misses evidence the score underestimates |
| DuoAttention | Head classification at load time | Up to 2.18× decode, 1.73× prefill | A misclassified retrieval head loses context |
| H2O / SnapKV / PyramidKV | Irreversible eviction | 8×+ memory, direct decode bandwidth | Drops the needle; poisons cached prefixes |
| FP8 KV | ~0–0.5% quality | 2× cache bytes, 2× decode headroom | Outlier sensitivity at long context |
Further reading
- Liu et al., "SnapKV: LLM Knows What You Are Looking For Before Generation" (2024).
- Xiao et al., "Efficient Streaming Language Models with Attention Sinks" (ICLR 2024).
- Zhang et al., "H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models" (NeurIPS 2023).
- Cai et al., "PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling" (2024).
- Tang et al., "Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference" (ICML 2024).
- Xiao et al., "DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads" (2024).
- Liu et al., "RULER: What's the Real Context Size of Your Long-Context Language Models?" (2024).
- Dao et al., "FlashAttention" and Liu et al., "Ring Attention with Blockwise Transformers" (2023).