Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The autoregressive loop as a systems problem

One token at a time

Part 11 of the training guide's inference chapter introduced the KV cache and the prefill/decode split. That page answers what happens. This series answers how much: how many bytes move, how many FLOPs exist to hide them, and where a real machine actually spends its time.

A decoder-only transformer emits token t+1 by running a forward pass over all tokens so far, reading the final logits, sampling, and appending. The load-bearing phrase is all tokens so far: attention needs the keys and values of every earlier position, so an engine caches them and computes only the newest position each step. That cache is why generation is linear in output length instead of quadratic. It is also, as this page shows, why generation is slow.

Strip a decode step to its bones and three things happen. Every weight in the model is read from HBM to multiply the single new activation vector. Every resident sequence's KV cache is read for attention. A tiny number of FLOPs is issued — roughly 2P per token for P parameters. The step time is therefore set by bytes, not arithmetic:

$$T_{\text{step}} \;\approx\; \frac{P \cdot b_p \;+\; \text{KV}_{\text{read}}}{\text{BW}} \;+\; T_{\text{overhead}}$$

where P is parameters, bp is bytes per parameter, and BW is HBM bandwidth. For a 70B model in bf16 that numerator starts at 140 GB, before any KV. At H100 bandwidth (3.35 TB/s) the floor is about 42 ms per token — roughly 24 tokens/s for one stream, no matter how fast the tensor cores are. This is the memory wall: generation is a bandwidth problem wearing a compute costume.

💡 Carry one number through the series: a decode step is worth on the order of 1–2 FLOPs per byte of memory traffic. Every optimization from here — batching, quantization, speculation, disaggregation — is a way of changing that ratio or the traffic it multiplies.

2

Prefill and decode are two different programs

Same weights, opposite bottlenecks

Prefill consumes the whole prompt at once. It is one large matrix multiplication over Pprompt tokens: weights are read once and reused Pprompt times, so the arithmetic intensity is enormous and the GPU runs compute-bound near its tensor-core ceiling. Decode consumes one token per sequence per step. Weights are read once and reused once, so intensity collapses and the GPU stalls on HBM. The GQA and MLA mechanisms shrink the KV term — the second term above — but they cannot touch the weight term, which is why they buy concurrency rather than raw single-stream speed.

The practical consequence is that one request is two workloads with a phase change in the middle, and the two phases want different hardware. Watch the shape of the work: prefill fills a grid of positions in a single sweep; decode adds positions one at a time.

2,000 positions. Prefill processes all of them in one pass; decode adds exactly one per tick. This is why TTFT and TPOT are measured separately — they are different machines.

3

Arithmetic intensity

FLOPs per byte decides everything

Arithmetic intensity is FLOPs divided by bytes moved from HBM. For a decode step, FLOPs are 2P·B for batch B and bytes are roughly P·bp for the weights; the P cancels, leaving intensity that depends only on precision and batch:

$$\text{AI}_{\text{decode}} \;\approx\; \frac{2B}{b_p}\text{ FLOPs/byte},\qquad \text{AI}_{\text{prefill}} \;\approx\; \frac{2P_{\text{prompt}}}{b_p}\text{ FLOPs/byte}$$

In bf16, single-stream decode sits at about 1 FLOP/byte. A 2,048-token prefill sits near 2,048 FLOPs/byte — a gap of three orders of magnitude inside the same call. Raising batch is the one knob that moves decode intensity without changing the model: each extra sequence in the batch reuses the same weight read for another set of FLOPs. The KV cache, though, grows with batch and context, so it eventually reclaims the byte budget. Move the sliders below to watch weights dominate at low batch and KV take over at high batch.

Bytes read per decode step, split into weights (accent) and KV cache (magenta). Shared linear scale, so the bars show absolute traffic.

4

The roofline

Two ceilings, one knee

A roofline plots attainable throughput against arithmetic intensity. On the left of the knee the bandwidth line AI · BW binds; on the right the flat compute ceiling peak binds. The knee sits at peak / BW, the ridge point. For an H100 that ridge is 295 FLOPs/byte (989 TFLOP/s ÷ 3.35 TB/s); for a B200 it is about 281 (2,250 ÷ 8). Faster silicon moves the ceilings but not the shape, and decode at batch 1 lives far to the left of the knee on every accelerator ever built.

Prefill and decode are two points on this same curve, and batching slides the decode point rightward. Drag either point: the decode point changes batch, the prefill point changes prompt length. The verdict under the canvas says which ceiling owns each point.

Log–log axes. Drag the prefill point horizontally to change prompt length; drag the decode point to change batch.

5

Batching as the way up it

Reuse the weight read

If one decode step must read 140 GB of weights, the obvious move is to amortize that read over more work. Batching B sequences does exactly that: the same weight sweep produces B tokens. Aggregate throughput therefore rises almost linearly with batch while the step stays memory-bound — and then stops rising once the batch is large enough to push the step over the ridge into compute-bound territory. From that point on, more batch adds KV traffic and queueing delay without adding throughput. The long-context chapter owns what happens when the KV term is the dominant one.

The curve below is the whole argument in one picture. It uses the same bytes-vs-FLOPs model as the rest of the page and marks the batch where decode crosses the ridge.

Aggregate tokens/s across a batch of resident sequences, computed with the step-time model. The dashed line is the ridge crossing.

💡 The batching rule: throughput scales with batch until the step becomes compute-bound, then flattens. Part 6 turns this static picture into continuous batching, where the batch is refilled every iteration rather than fixed per group.
6

MFU vs MBU

Two utilizations, one honest dashboard

Model FLOPs utilization (MFU) is achieved FLOPs over peak FLOPs; memory-bandwidth utilization (MBU) is achieved bytes/s over peak bandwidth. A memory-bound workload cannot have both. At batch 1 the decode step above runs at roughly 100% MBU and well under 1% MFU — the HBM is saturated while the tensor cores idle. Reporting MFU alone therefore makes a perfectly tuned decode server look broken, and chasing MFU leads straight into the trap below.

Toggle the two levers on a single decode step and watch which one moves the needle. Doubling peak FLOPs changes step time by a rounding error; doubling HBM bandwidth nearly halves it. The readout reports the change in aggregate tokens/s and the live MFU/MBU split.

Aggregate tokens/s for the selected batch under four hardware assumptions.

⚠️ The "twice the FLOPs" trap: at low batch an accelerator with double the tensor-core peak buys a fraction of a percent on decode, while double the memory bandwidth buys nearly 2×. Match the chip to the phase; prefill wants FLOPs, decode wants bytes. Part 2 maps the accelerator zoo around exactly this split.
7

Map of the series

Where each remaining part fits

Everything after this page is a lever on one of the three terms above: the weight bytes (quantization, parallelism, MoE), the KV bytes (paging, prefix caching, GQA/MLA, long context), or the effective batch and its scheduling (continuous batching, schedulers, disaggregation). The glossary collects every constant cited here.

Cheat sheet

QuantityDecodePrefill
Arithmetic intensity (bf16)≈ batch, ~1 FLOP/byte at batch 1≈ prompt tokens, thousands of FLOPs/byte
BottleneckHBM bandwidthTensor-core FLOPs
Bytes per stepP·bp + KV(batch, context)P·bp once, reused Pprompt×
Ridge pointpeak ÷ bandwidth ≈ 295 FLOPs/byte on H100, ≈ 281 on B200
Right utilization metricMBU (want ~100%), MFU is misleadingMFU (want high), bandwidth rarely binds
Main leverBatch to amortize the weight read, then quantizeFaster tensor cores, chunking to protect ITL

Further reading

?

Check your understanding

0/5 answered