Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Why the cache exists, and why it dominates serving

From a training-time trick to the binding constraint

Autoregressive decoding needs every earlier token's key and value vectors to attend to them, so recomputing them at each step would make generation quadratic for no reason. Instead the K and V tensors are cached once and reused, and each step computes only the newest token's — the LLM Training guide introduced the mechanism. What that page does not quantify is the serving consequence: the cache is per-sequence, grows linearly with context, and must be resident for every concurrent request, so it competes directly with the weights for HBM.

At production scale the cache stops being an optimization and becomes the system. DeepSeek reported that over one 24-hour window, 342 billion of 608 billion input tokens — 56.3% — were served from a cache rather than recomputed, which is why the rest of this series treats KV as a first-class, paged, routed, quantized resource. Before any of that machinery helps, you have to know how many bytes one token costs.

💡 Where this is going: the cache is the reason Part 5 needs paged memory, Part 8 needs prefix caching, and Part 13 needs to ship KV across a network. Every one of those costs is measured in bytes per token, so get the formula right first.
2

The formula, term by term

One multiplication decides the whole budget

For standard attention the cache holds a key and a value tensor for every KV head of every layer, at the model's head dimension, for every token in the sequence, at the chosen element width, for every sequence in the batch.

$$\text{KV bytes} = 2 \times L \times n_{kv} \times d_{head} \times T \times b \times \text{batch}$$

2 · K and V  |  L · layers  |  n_kv · KV heads  |  d_head · head dim  |  T · tokens  |  b · bytes per element  |  batch · concurrent sequences

Every term is a lever. The 2 is fixed by attention. L and d_head are architectural and expensive to change. n_kv is the lever GQA and MQA pull. T is the context you agreed to serve. b is the lever KV quantization pulls. Batch is the consequence — the number you are solving for, not a knob. Change any term below and watch bytes per token move.

Total cache for the whole batch, and how it splits.

3

Bytes per token for real models

The configs everyone actually runs

The table below computes the formula from each model's published config at FP16. Two things jump out. First, a 405B dense model with 8 KV heads costs about the same per token as a 70B — the cache is a function of layers and heads, not parameter count. Second, the MoE models look enormous by weights but cheap by cache, because their KV cost is set by layer count and attention heads while most of their parameters live in experts that never touch the cache.

ModelSchemeLayersKV headsBytes/token (FP16)Cache at full context

💡 Read the third column, not the first: Llama 3.1 405B and 70B share 8 KV heads and differ by 1.6× in layers, so their caches differ by roughly that much — not by the 5.7× that separates their parameter counts.
4

MHA, GQA, MQA

Trading a little quality for a lot of cache

Multi-head attention caches a K and V for every query head. Grouped-query attention gives several query heads a single shared K/V, and multi-query attention shares one K/V across all of them. The mechanism — several query heads reading the same subspace — is drawn in the architecture page; what this page owns is the accounting. GQA with 8 groups is an 8× cache reduction at negligible quality cost and is now the default in Llama 3.1, Mistral and gpt-oss. MQA is a 64× reduction but can hurt quality, which is why frontier models rarely ship it.

Each line is a query head reading from its KV head.

5

MLA's latent

Cache the compression, not the heads

Multi-head latent attention takes a different route: instead of caching per-head K and V, it caches a single low-rank latent per layer and reconstructs the per-head keys and values on the fly during attention. At DeepSeek-V3's geometry that is a 512-element latent plus a 64-element decoupled RoPE key — 576 elements per layer per token, or about 70 KB per token in FP16 across 61 layers. The MHA equivalent at the same depth and 128 heads would be roughly 4.0 MB/token, so the compression is on the order of 57×, achieved by moving work from memory into compute. The cache stops being a per-head array and becomes one compact vector per layer, which is exactly what makes long-context serving tractable for these models.

$$\text{MLA bytes/token} = L \times (d_c + d_{rope}) \times b = 61 \times (512 + 64) \times 2 \approx 70\,\text{KB}$$
6

Sliding windows and hybrid layers

Making the cache flat instead of linear

A sliding-window layer attends only to the most recent W tokens, so its cache stops growing once the sequence passes W — peak memory becomes bounded rather than linear in context. Mistral 7B ships a 4,096-token window natively; gpt-oss alternates sliding and full layers, so only a fraction of layers grow with the sequence. State-space and hybrid SSM layers (Mamba-style) go further and keep a constant-size recurrent state instead of a per-token cache at all, at the cost of losing exact long-range retrieval. The trade is always the same: you buy bounded memory with a bounded attention horizon.

KV per sequence as the sequence grows. Play to sweep the length.

7

The weights / KV / activations budget

Concurrency is what is left after everything else

A node's HBM has to hold three things at once: the sharded weights, one KV cache per resident sequence, and a working set of activations. Weights are a fixed cost paid once. Activations are small per sequence in decode. The KV cache is the term that scales with load, so it is the term that converts HBM into concurrency: the maximum number of sequences a GPU can hold is whatever capacity remains after weights, divided by the per-sequence cache. Halving cache bytes per token doubles concurrency; doubling context halves it.

Total resident GB versus concurrency; the dashed line is HBM capacity.

💡 Which lever moves it most: at long context the KV term dwarfs everything, so KV quantization beats a bigger GPU almost every time. At short context the weights dominate and only capacity or weight precision helps. The readout names the winner for the current setting.
8

KV quantization

The cheapest bandwidth win in serving

Because the cache is read once per token per sequence, quantizing it to FP8 halves both the memory it occupies and the bytes decode has to stream — often a near-free win, since FP8 KV typically costs little quality when keys and values are scaled per group. INT4 KV is more aggressive: it can cut the cache four-fold, but keys are unusually sensitive to outliers, so quality damage is far more visible than the nearly-lossless INT4 weights that Part 10 discusses. The practical ranking is FP16 baseline, FP8 broadly safe, INT4 only with per-group scales and task-level evaluation. Every byte removed here is a byte of concurrency added back, which is why KV precision is the first knob a serving engineer reaches for.

⚠️ The asymmetry: weight quantization changes what the model knows; KV quantization changes what it can still attend to. Damage from INT4 KV often shows up as long-range recall failures that a short benchmark never sees.
9

From bytes to concurrency

The one number to carry into capacity planning

Putting the pieces together, the concurrency a single GPU supports is:

$$N_{\max} = \left\lfloor \frac{\text{HBM}_{\text{usable}} - \text{weights}}{\text{bytes/token} \times T + \text{activations per sequence}} \right\rfloor$$

Everything else in this series — paging, prefix sharing, admission control — is about making the effective numerator larger or the denominator smaller. Paging removes fragmentation so the numerator is actually usable; prefix caching makes shared tokens in the denominator cost nothing; quantization shrinks bytes/token; sliding windows cap T. The rest of the guide is the machinery around this one division. Next, Part 5 shows how the numerator is managed in fixed-size blocks instead of one contiguous reservation.

📌 Cross-link: the weights term here is the inference-time cost; its training-time twin (optimizer state, gradients, master weights) is worked through in the scaling page's memory accounting. Same parameters, roughly 8× the bytes.
✓

Cheat sheet

The cache in six lines

QuantityFormula / valueLever
Bytes per token2 · L · nkv · dhead · bThe master equation
MHA vs GQA vs MQA2.62 MB → 0.33 MB → 0.04 MB per token (FP16, 70B geometry)Share K/V across query heads
MLAL · (d_c + d_rope) · b ≈ 70 KB/token (61 layers, 576)Cache a latent
Sliding windowCache flat at W instead of linear in TBound the attention horizon
KV FP8Half the bytes, near-free qualityElement width b
Max concurrency(HBM − weights) ÷ (bytes/token · T + activations)The planning number
📚

Further reading

References

?

Check your understanding

0/5 answered