Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A format is just a bit budget

Sign, exponent, mantissa

The training guide's quantization demo established the basic size-versus-error trade by rounding a synthetic weight distribution at 16, 8 and 4 bits. That is the right first intuition and we will not repeat it. The serving question is narrower and harder: given 16, 8 or 4 bits, how are they spent, and where does the scale that maps them back to real numbers live? A format is a fixed contract — so many bits of sign, so many of exponent, so many of mantissa — and every downstream number (bytes per token, concurrent sequences, tokens per second, and a small quality tax) is a consequence of that contract.

Pick a format and a probe value below. The strip shows how the bits are allocated, the number line shows which real values the format can actually represent near the probe, and the readout reports the round-trip error you would introduce by storing the probe in that format.

Ticks are representable values in the window around the probe; the marker is the probe, the ring is the nearest representable value.

sign exponent mantissa / magnitude
2

Sixteen bits, eight bits, and where the scale lives

Width is only half the decision

FP16 and BF16 are both 16 bits and are not interchangeable. FP16 spends 5 bits on the exponent and 10 on the mantissa; BF16 spends 8 and 7. That gives BF16 the same exponent range as FP32 — roughly 10±38 — and FP16 more mantissa near 1.0. In practice BF16 is preferred for weights and activations because a wide exponent tolerates the occasional large value without saturating, and the 8-bit mantissa is plenty once you accept a relative step of about 2-8 ≈ 0.4%.

FP8 comes in two shapes for two jobs. E4M3 (4 exponent bits, 3 mantissa) keeps precision up to a maximum finite value of 448 and is the format Hopper and Blackwell tensor cores accelerate natively; DeepSeek-V3 trained and served in E4M3 with per-block scaling. E5M2 trades two mantissa bits for a much larger range and is used where values can spike, classically in gradients. Neither is a drop-in for the 16-bit baseline: at 3 mantissa bits the relative step is 2-3 = 12.5%, which is why FP8 only works with a scale.

And that is the part the bit-width conversation always skips. Quantization error is set by the range you pack into a fixed number of levels, and the range is chosen per scale. One scale per tensor means a single large value dictates the step for every other value. One scale per output channel is strictly better. One scale per group of 128 weights along the input dimension is better again, at the cost of storing more scales. The readout above uses a fixed per-tensor range; the next demo makes the choice explicit.

3

Outliers: the tensor you actually have to quantize

SmoothQuant and the migration of scale

Weights are well behaved: roughly Gaussian, similar magnitude across channels. Activations are not. A handful of input channels — often the same few across layers — carry values one to two orders of magnitude larger than the rest (the "massive activations" that make W8A8 harder than weight-only). Quantize activations per-tensor and one outlier channel sets a step so coarse that every ordinary channel is destroyed. The fix, from SmoothQuant, is an algebraic identity: you may multiply a column of the weights by a per-channel factor and divide the corresponding activation channel by the same factor without changing the product. Choose the factor to move the activation's dynamic range into the weights, which tolerate it.

The demo builds a seeded Gaussian weight matrix and an activation matrix with three deliberate outlier channels, then quantizes both under your chosen format and granularity. Slide alpha to move dynamic range out of the activations and into the weights: the activation profile flattens monotonically, but the output error has a shallow optimum in between — too little alpha leaves the outliers in the activations, while pushing all of it into the weights makes those weights harder to quantize. Production SmoothQuant sits near α = 0.5.

Per-input-channel magnitude on a log₂ scale. Spikes are outliers; the accent line is after smoothing.

4

INT4 weight-only: GPTQ and AWQ

The safest 4 bits

The most common production configuration is deliberately asymmetric: weights in 4 bits, activations left in FP16. This sidesteps the outlier problem entirely, because nothing about the activations changes — only the weight matrices are stored coarsely, and each is dequantized to FP16 immediately before the matmul. GPTQ does this with layer-wise second-order error compensation: it quantizes one column at a time and updates the remaining weights to cancel the error it just introduced, reaching 3–4 bits with negligible degradation and a reported ~3.25× end-to-end speedup on an A100. AWQ starts from the observation that about 1% of channels matter far more than the rest and protects them by scaling rather than by keeping them at higher precision, which is cheaper and tends to be the safer default for instruction-tuned models. Group size matters more than the algorithm in practice: 128 is standard, and going below 64 buys little quality for a lot of scale overhead.

The catch is the one section 7 makes visual: weights in 4 bits shrink the bytes you read, but the tensor cores still run the matmul at FP16 precision unless the activations are quantized too. At batch size 1, where decode is purely memory-bound, that is almost a pure win. At high batch, where the step becomes compute-bound, it is worth nothing.

5

Block formats: MXFP4 and NVFP4

Shared exponents instead of shared scales

There is a third way to spend the scale bits. Instead of a floating-point scale per block, microscaling (MX) formats store a single power-of-two exponent shared by a block of values, and each value is a small set of mantissa bits — for MXFP4, an E2M1 element with representable magnitudes {0, 0.5, 1, 1.5, 2, 3, 4, 6}. The block size is fixed at 32 and the exponent is shared by all 32 values, so the format is dequantize-friendly: hardware can reconstruct the block cheaply, and there is no per-tensor range to blow up. NVFP4 is NVIDIA's Blackwell variant, adding a per-block FP8 scale on top of the per-value E2M1 (a two-level scaling scheme, block 16) so a block's values need not be exact powers of two. gpt-oss-120b ships its weights in MXFP4; Blackwell adds native FP4 tensor cores, which is why B200 advertises ~9,000 dense FP4 TFLOP/s (the H100 column in the numbers table has no FP4 row at all).

Block formats move the quality question from "how many bits" to "how many bits, and which layers stay wide". Community evaluations consistently find that the damage concentrates in attention projections, router weights and embeddings, and that keeping those at BF16 while the feed-forward blocks go to 4 bits is worth more than any choice of 4-bit algorithm.

6

KV-cache quantization is a separate decision

A second, larger tensor

Weights are static and read once per token; the KV cache is written once and then re-read on every single decode step, and it grows with context and concurrency. It is a different tensor with a different access pattern and it deserves its own format decision — you can ship FP8 weights and FP16 KV, or BF16 weights and INT8 KV, independently. The lever is enormous because KV bytes are what cap how many sequences fit in a fixed memory pool. Following Part 4, a Llama-3.1-70B-class model with GQA caches 327,680 bytes per token at FP16 — 0.33 MB per token, before any batching. Halving the element width doubles the number of concurrent long-context sequences at the same memory budget.

Maximum resident sequences at the chosen context length, for a fixed KV memory budget.

KV quantization is harder to get right than weight quantization, because the cache is per-request and every error is re-read thousands of times. Per-channel keys and per-token values are the common pattern, and most engines keep the first and last few layers, or the attention sinks, at higher precision. Treat it as a knob to test against your own evaluation, not a free 2×.

7

The win is bandwidth, not FLOPs

Why the speedup evaporates at batch

This is the counter-intuitive result worth internalizing. Quantization shrinks the number of bytes a decode step must read. It does not, by itself, reduce the number of FLOPs the matmul performs: a 4-bit weight is dequantized to FP16 before the multiply, so the tensor core still does FP16 work unless the activation is quantized too. At batch size 1 every decode step is memory-bound (arithmetic intensity around 1–2 FLOPs/byte) and the step time is almost exactly bytes over bandwidth — so 4-bit weights look like a 4× win. As batch grows, the same weight read is amortized over more tokens and the step becomes compute-bound; the int4 kernel must still do all the FP16 FLOPs and the speedup collapses toward 1×. The fix is asymmetric again: quantize the activations too (W4A8, 4-bit weights with FP8 activations) so the matmul runs on the FP8 tensor cores, which keeps roughly a 2× even in the compute-bound regime. The plot below uses the shared step-time model: x-axis is batch, y-axis is speedup over the FP16 baseline at the same batch.

Solid marks show the selected batch on each curve; the dashed line is parity with BF16.

8

Calibration and measuring the damage

The accuracy bill

Per-channel and per-group scaling need to know the range of each channel, and for activations you cannot know it without running data through the model. That is calibration: a few hundred sequences, ideally drawn from the same distribution as production traffic, run through the model to record activation ranges. It is why "post-training quantization" is not actually post-training — the ranges are data-dependent, and a model quantized on English Wikipedia will behave differently on code or on a low-resource language. Calibration set size and composition are as much a hyperparameter as group size.

The table and bars below are the accuracy bill as reported in the primary papers. Read the deltas as orders of magnitude, not commitments.

⚠️ Every delta below is model and task dependent. A 1–3% average can hide a 20% collapse on a specific reasoning or code benchmark, and weight-only INT4 is usually far more benign than INT4 activations. Quantization must be evaluated on your own task suite and your own traffic; the numbers here are priors, not guarantees.
FormatBitsTypical deltaWhat it is for

✓

Cheat sheet: the decision table

Pick a format

SituationStart withWatch out for
You want the safest large win, single-stream latency mattersINT4 weight-only (AWQ), group 128, activations FP16Gains fade above batch ~128; measure at your real batch, not at batch 1
You serve high batch and want throughput, not just latencyW4A8, or FP8 E4M3 per-block on native hardwareNeeds activation calibration and outlier handling
You are on Hopper/Blackwell and want near-BF16 qualityFP8 E4M3 with per-block (128×128) scalesKeep the router, embeddings and attention output projections wide if quality moves
You need maximum memory saving and have a task suite to test againstMXFP4 / NVFP4 block formatsSensitive layers dominate the damage; block size and per-layer precision matter more than the algorithm
Concurrency is capped by the KV cache, not by weightsINT8 or FP8 KV, keeping sinks and boundary layers at FP16Every error is re-read thousands of times — evaluate, do not assume
You have no evaluation harness yetDo not quantize below BF16The accuracy bill is unmeasurable until you can measure it
📚

Further reading

References

?

Check your understanding

0/5 answered