Quantization and numerics
Every bit you drop is memory bandwidth you do not pay for on every token — and a small numerical error you may or may not notice. This part starts at the bit pattern, shows what an activation outlier does to it, and ends with a decision table you can put in a runbook.
A format is just a bit budget
Sign, exponent, mantissa
The training guide's quantization demo established the basic size-versus-error trade by rounding a synthetic weight distribution at 16, 8 and 4 bits. That is the right first intuition and we will not repeat it. The serving question is narrower and harder: given 16, 8 or 4 bits, how are they spent, and where does the scale that maps them back to real numbers live? A format is a fixed contract — so many bits of sign, so many of exponent, so many of mantissa — and every downstream number (bytes per token, concurrent sequences, tokens per second, and a small quality tax) is a consequence of that contract.
Pick a format and a probe value below. The strip shows how the bits are allocated, the number line shows which real values the format can actually represent near the probe, and the readout reports the round-trip error you would introduce by storing the probe in that format.
Ticks are representable values in the window around the probe; the marker is the probe, the ring is the nearest representable value.
Sixteen bits, eight bits, and where the scale lives
Width is only half the decision
FP16 and BF16 are both 16 bits and are not interchangeable. FP16 spends 5 bits on the exponent and 10 on the mantissa; BF16 spends 8 and 7. That gives BF16 the same exponent range as FP32 — roughly 10±38 — and FP16 more mantissa near 1.0. In practice BF16 is preferred for weights and activations because a wide exponent tolerates the occasional large value without saturating, and the 8-bit mantissa is plenty once you accept a relative step of about 2-8 ≈ 0.4%.
FP8 comes in two shapes for two jobs. E4M3 (4 exponent bits, 3 mantissa) keeps precision up to a maximum finite value of 448 and is the format Hopper and Blackwell tensor cores accelerate natively; DeepSeek-V3 trained and served in E4M3 with per-block scaling. E5M2 trades two mantissa bits for a much larger range and is used where values can spike, classically in gradients. Neither is a drop-in for the 16-bit baseline: at 3 mantissa bits the relative step is 2-3 = 12.5%, which is why FP8 only works with a scale.
And that is the part the bit-width conversation always skips. Quantization error is set by the range you pack into a fixed number of levels, and the range is chosen per scale. One scale per tensor means a single large value dictates the step for every other value. One scale per output channel is strictly better. One scale per group of 128 weights along the input dimension is better again, at the cost of storing more scales. The readout above uses a fixed per-tensor range; the next demo makes the choice explicit.
Outliers: the tensor you actually have to quantize
SmoothQuant and the migration of scale
Weights are well behaved: roughly Gaussian, similar magnitude across channels. Activations are not. A handful of input channels — often the same few across layers — carry values one to two orders of magnitude larger than the rest (the "massive activations" that make W8A8 harder than weight-only). Quantize activations per-tensor and one outlier channel sets a step so coarse that every ordinary channel is destroyed. The fix, from SmoothQuant, is an algebraic identity: you may multiply a column of the weights by a per-channel factor and divide the corresponding activation channel by the same factor without changing the product. Choose the factor to move the activation's dynamic range into the weights, which tolerate it.
The demo builds a seeded Gaussian weight matrix and an activation matrix with three deliberate outlier channels, then quantizes both under your chosen format and granularity. Slide alpha to move dynamic range out of the activations and into the weights: the activation profile flattens monotonically, but the output error has a shallow optimum in between — too little alpha leaves the outliers in the activations, while pushing all of it into the weights makes those weights harder to quantize. Production SmoothQuant sits near α = 0.5.
Per-input-channel magnitude on a log₂ scale. Spikes are outliers; the accent line is after smoothing.
INT4 weight-only: GPTQ and AWQ
The safest 4 bits
The most common production configuration is deliberately asymmetric: weights in 4 bits, activations left in FP16. This sidesteps the outlier problem entirely, because nothing about the activations changes — only the weight matrices are stored coarsely, and each is dequantized to FP16 immediately before the matmul. GPTQ does this with layer-wise second-order error compensation: it quantizes one column at a time and updates the remaining weights to cancel the error it just introduced, reaching 3–4 bits with negligible degradation and a reported ~3.25× end-to-end speedup on an A100. AWQ starts from the observation that about 1% of channels matter far more than the rest and protects them by scaling rather than by keeping them at higher precision, which is cheaper and tends to be the safer default for instruction-tuned models. Group size matters more than the algorithm in practice: 128 is standard, and going below 64 buys little quality for a lot of scale overhead.
The catch is the one section 7 makes visual: weights in 4 bits shrink the bytes you read, but the tensor cores still run the matmul at FP16 precision unless the activations are quantized too. At batch size 1, where decode is purely memory-bound, that is almost a pure win. At high batch, where the step becomes compute-bound, it is worth nothing.
Block formats: MXFP4 and NVFP4
Shared exponents instead of shared scales
There is a third way to spend the scale bits. Instead of a floating-point scale per block, microscaling (MX) formats store a single power-of-two exponent shared by a block of values, and each value is a small set of mantissa bits — for MXFP4, an E2M1 element with representable magnitudes {0, 0.5, 1, 1.5, 2, 3, 4, 6}. The block size is fixed at 32 and the exponent is shared by all 32 values, so the format is dequantize-friendly: hardware can reconstruct the block cheaply, and there is no per-tensor range to blow up. NVFP4 is NVIDIA's Blackwell variant, adding a per-block FP8 scale on top of the per-value E2M1 (a two-level scaling scheme, block 16) so a block's values need not be exact powers of two. gpt-oss-120b ships its weights in MXFP4; Blackwell adds native FP4 tensor cores, which is why B200 advertises ~9,000 dense FP4 TFLOP/s (the H100 column in the numbers table has no FP4 row at all).
Block formats move the quality question from "how many bits" to "how many bits, and which layers stay wide". Community evaluations consistently find that the damage concentrates in attention projections, router weights and embeddings, and that keeping those at BF16 while the feed-forward blocks go to 4 bits is worth more than any choice of 4-bit algorithm.
KV-cache quantization is a separate decision
A second, larger tensor
Weights are static and read once per token; the KV cache is written once and then re-read on every single decode step, and it grows with context and concurrency. It is a different tensor with a different access pattern and it deserves its own format decision — you can ship FP8 weights and FP16 KV, or BF16 weights and INT8 KV, independently. The lever is enormous because KV bytes are what cap how many sequences fit in a fixed memory pool. Following Part 4, a Llama-3.1-70B-class model with GQA caches 327,680 bytes per token at FP16 — 0.33 MB per token, before any batching. Halving the element width doubles the number of concurrent long-context sequences at the same memory budget.
Maximum resident sequences at the chosen context length, for a fixed KV memory budget.
KV quantization is harder to get right than weight quantization, because the cache is per-request and every error is re-read thousands of times. Per-channel keys and per-token values are the common pattern, and most engines keep the first and last few layers, or the attention sinks, at higher precision. Treat it as a knob to test against your own evaluation, not a free 2×.
The win is bandwidth, not FLOPs
Why the speedup evaporates at batch
This is the counter-intuitive result worth internalizing. Quantization shrinks the number of bytes a decode step must read. It does not, by itself, reduce the number of FLOPs the matmul performs: a 4-bit weight is dequantized to FP16 before the multiply, so the tensor core still does FP16 work unless the activation is quantized too. At batch size 1 every decode step is memory-bound (arithmetic intensity around 1–2 FLOPs/byte) and the step time is almost exactly bytes over bandwidth — so 4-bit weights look like a 4× win. As batch grows, the same weight read is amortized over more tokens and the step becomes compute-bound; the int4 kernel must still do all the FP16 FLOPs and the speedup collapses toward 1×. The fix is asymmetric again: quantize the activations too (W4A8, 4-bit weights with FP8 activations) so the matmul runs on the FP8 tensor cores, which keeps roughly a 2× even in the compute-bound regime. The plot below uses the shared step-time model: x-axis is batch, y-axis is speedup over the FP16 baseline at the same batch.
Solid marks show the selected batch on each curve; the dashed line is parity with BF16.
Calibration and measuring the damage
The accuracy bill
Per-channel and per-group scaling need to know the range of each channel, and for activations you cannot know it without running data through the model. That is calibration: a few hundred sequences, ideally drawn from the same distribution as production traffic, run through the model to record activation ranges. It is why "post-training quantization" is not actually post-training — the ranges are data-dependent, and a model quantized on English Wikipedia will behave differently on code or on a low-resource language. Calibration set size and composition are as much a hyperparameter as group size.
The table and bars below are the accuracy bill as reported in the primary papers. Read the deltas as orders of magnitude, not commitments.
| Format | Bits | Typical delta | What it is for |
|---|
Cheat sheet: the decision table
Pick a format
| Situation | Start with | Watch out for |
|---|---|---|
| You want the safest large win, single-stream latency matters | INT4 weight-only (AWQ), group 128, activations FP16 | Gains fade above batch ~128; measure at your real batch, not at batch 1 |
| You serve high batch and want throughput, not just latency | W4A8, or FP8 E4M3 per-block on native hardware | Needs activation calibration and outlier handling |
| You are on Hopper/Blackwell and want near-BF16 quality | FP8 E4M3 with per-block (128×128) scales | Keep the router, embeddings and attention output projections wide if quality moves |
| You need maximum memory saving and have a task suite to test against | MXFP4 / NVFP4 block formats | Sensitive layers dominate the damage; block size and per-layer precision matter more than the algorithm |
| Concurrency is capped by the KV cache, not by weights | INT8 or FP8 KV, keeping sinks and boundary layers at FP16 | Every error is re-read thousands of times — evaluate, do not assume |
| You have no evaluation harness yet | Do not quantize below BF16 | The accuracy bill is unmeasurable until you can measure it |
Further reading
References
- Frantar et al., GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, ICLR 2023.
- Lin et al., AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration, MLSys 2024.
- Xiao et al., SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models, ICML 2023.
- DeepSeek-AI, DeepSeek-V3 Technical Report, 2024 — FP8 E4M3 training and inference at scale.
- NVIDIA, NVFP4 explained, 2025 — microscopic 4-bit block formats and Blackwell FP4 tensor cores.
- Rouhani et al., Microscaling Data Formats for Deep Learning, 2023 — the MX specification.