The KV cache: sizing, GQA, MQA and MLA
The cache is the single largest allocation in a serving node once weights are fixed, and it is the thing that decides how many users a GPU can hold at once. This part turns it into arithmetic: one formula, run against real model configs, then attacked with every architectural lever — grouped heads, a learned latent, sliding windows — until the number moves by an order of magnitude.
Why the cache exists, and why it dominates serving
From a training-time trick to the binding constraint
Autoregressive decoding needs every earlier token's key and value vectors to attend to them, so recomputing them at each step would make generation quadratic for no reason. Instead the K and V tensors are cached once and reused, and each step computes only the newest token's — the LLM Training guide introduced the mechanism. What that page does not quantify is the serving consequence: the cache is per-sequence, grows linearly with context, and must be resident for every concurrent request, so it competes directly with the weights for HBM.
At production scale the cache stops being an optimization and becomes the system. DeepSeek reported that over one 24-hour window, 342 billion of 608 billion input tokens — 56.3% — were served from a cache rather than recomputed, which is why the rest of this series treats KV as a first-class, paged, routed, quantized resource. Before any of that machinery helps, you have to know how many bytes one token costs.
The formula, term by term
One multiplication decides the whole budget
For standard attention the cache holds a key and a value tensor for every KV head of every layer, at the model's head dimension, for every token in the sequence, at the chosen element width, for every sequence in the batch.
2 · K and V | L · layers | n_kv · KV heads | d_head · head dim | T · tokens | b · bytes per element | batch · concurrent sequences
Every term is a lever. The 2 is fixed by attention. L and d_head are architectural and expensive to change. n_kv is the lever GQA and MQA pull. T is the context you agreed to serve. b is the lever KV quantization pulls. Batch is the consequence — the number you are solving for, not a knob. Change any term below and watch bytes per token move.
Total cache for the whole batch, and how it splits.
Bytes per token for real models
The configs everyone actually runs
The table below computes the formula from each model's published config at FP16. Two things jump out. First, a 405B dense model with 8 KV heads costs about the same per token as a 70B — the cache is a function of layers and heads, not parameter count. Second, the MoE models look enormous by weights but cheap by cache, because their KV cost is set by layer count and attention heads while most of their parameters live in experts that never touch the cache.
| Model | Scheme | Layers | KV heads | Bytes/token (FP16) | Cache at full context |
|---|
MHA, GQA, MQA
Trading a little quality for a lot of cache
Multi-head attention caches a K and V for every query head. Grouped-query attention gives several query heads a single shared K/V, and multi-query attention shares one K/V across all of them. The mechanism — several query heads reading the same subspace — is drawn in the architecture page; what this page owns is the accounting. GQA with 8 groups is an 8× cache reduction at negligible quality cost and is now the default in Llama 3.1, Mistral and gpt-oss. MQA is a 64× reduction but can hurt quality, which is why frontier models rarely ship it.
Each line is a query head reading from its KV head.
MLA's latent
Cache the compression, not the heads
Multi-head latent attention takes a different route: instead of caching per-head K and V, it caches a single low-rank latent per layer and reconstructs the per-head keys and values on the fly during attention. At DeepSeek-V3's geometry that is a 512-element latent plus a 64-element decoupled RoPE key — 576 elements per layer per token, or about 70 KB per token in FP16 across 61 layers. The MHA equivalent at the same depth and 128 heads would be roughly 4.0 MB/token, so the compression is on the order of 57×, achieved by moving work from memory into compute. The cache stops being a per-head array and becomes one compact vector per layer, which is exactly what makes long-context serving tractable for these models.
Sliding windows and hybrid layers
Making the cache flat instead of linear
A sliding-window layer attends only to the most recent W tokens, so its cache stops growing once the sequence passes W — peak memory becomes bounded rather than linear in context. Mistral 7B ships a 4,096-token window natively; gpt-oss alternates sliding and full layers, so only a fraction of layers grow with the sequence. State-space and hybrid SSM layers (Mamba-style) go further and keep a constant-size recurrent state instead of a per-token cache at all, at the cost of losing exact long-range retrieval. The trade is always the same: you buy bounded memory with a bounded attention horizon.
KV per sequence as the sequence grows. Play to sweep the length.
The weights / KV / activations budget
Concurrency is what is left after everything else
A node's HBM has to hold three things at once: the sharded weights, one KV cache per resident sequence, and a working set of activations. Weights are a fixed cost paid once. Activations are small per sequence in decode. The KV cache is the term that scales with load, so it is the term that converts HBM into concurrency: the maximum number of sequences a GPU can hold is whatever capacity remains after weights, divided by the per-sequence cache. Halving cache bytes per token doubles concurrency; doubling context halves it.
Total resident GB versus concurrency; the dashed line is HBM capacity.
KV quantization
The cheapest bandwidth win in serving
Because the cache is read once per token per sequence, quantizing it to FP8 halves both the memory it occupies and the bytes decode has to stream — often a near-free win, since FP8 KV typically costs little quality when keys and values are scaled per group. INT4 KV is more aggressive: it can cut the cache four-fold, but keys are unusually sensitive to outliers, so quality damage is far more visible than the nearly-lossless INT4 weights that Part 10 discusses. The practical ranking is FP16 baseline, FP8 broadly safe, INT4 only with per-group scales and task-level evaluation. Every byte removed here is a byte of concurrency added back, which is why KV precision is the first knob a serving engineer reaches for.
From bytes to concurrency
The one number to carry into capacity planning
Putting the pieces together, the concurrency a single GPU supports is:
Everything else in this series — paging, prefix sharing, admission control — is about making the effective numerator larger or the denominator smaller. Paging removes fragmentation so the numerator is actually usable; prefix caching makes shared tokens in the denominator cost nothing; quantization shrinks bytes/token; sliding windows cap T. The rest of the guide is the machinery around this one division. Next, Part 5 shows how the numerator is managed in fixed-size blocks instead of one contiguous reservation.
Cheat sheet
The cache in six lines
| Quantity | Formula / value | Lever |
|---|---|---|
| Bytes per token | 2 · L · nkv · dhead · b | The master equation |
| MHA vs GQA vs MQA | 2.62 MB → 0.33 MB → 0.04 MB per token (FP16, 70B geometry) | Share K/V across query heads |
| MLA | L · (d_c + d_rope) · b ≈ 70 KB/token (61 layers, 576) | Cache a latent |
| Sliding window | Cache flat at W instead of linear in T | Bound the attention horizon |
| KV FP8 | Half the bytes, near-free quality | Element width b |
| Max concurrency | (HBM − weights) ÷ (bytes/token · T + activations) | The planning number |
Further reading
References
- Ainslie et al., "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints" (EMNLP 2023).
- Shazeer, "Fast Transformer Decoding: One Write-Head is All You Need" (2019) — multi-query attention.
- DeepSeek-AI, "DeepSeek-V3 Technical Report" (arXiv:2412.19437) — multi-head latent attention.
- Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention" (SOSP 2023) — why cache bytes become blocks.