Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The end-to-end path across process boundaries

One request, six handoffs, at least two processes

The training guide's serving page stops at the GPU: prefill then decode, batched continuously. A real engine has to move text across process boundaries before and after that. A gateway terminates HTTP, authenticates and enforces quotas; a tokenizer process converts text to ids; the scheduler process — vLLM V1 calls it EngineCore — owns the model and decides what runs each iteration; the GPU worker executes kernels; the detokenizer converts ids back; and a streaming layer flushes deltas to the client. The reason this matters at serving scale is that every handoff adds latency the GPU cannot hide, and the naive implementation serialises them. Production engines pipeline the stages and keep the model loop in a separate process precisely so a slow JSON parse or a blocking socket write cannot stall a decode iteration.

Click any stage in the strip. The batch slider at the bottom is the thread that ties the whole part together: the fraction of a step spent launching kernels on the CPU collapses as batch size grows, which is why the expensive machinery in the second half of this page only starts to pay at small batch.

Click a stage; drag the batch slider to move the CPU/GPU split.

2

The tokenizer, the detokenizer and the streaming path

The bytes you actually stream are not tokens

Tokenization is CPU work measured in tens of microseconds for a normal prompt — irrelevant next to a prefill — but it is on the critical path, and a pathological prompt (long, emoji-heavy, or byte-fallback-heavy) can cost tens of milliseconds. The output side is subtler. A token is a byte sequence, not necessarily valid UTF-8, so a naive detokenizer that emits each token as it is sampled can split a multi-byte character across two SSE events and produce mojibake. Correct streaming detokenizers buffer until a token boundary is a valid character boundary, apply stop-string matching across deltas, and only then flush. That buffering is why "streaming" in a well-built engine is not token-by-token: it is the smallest safe text delta the client can append.

💡 Where the time goes: for a 2,000-token prompt the tokenizer is well under 1% of TTFT, so engines accept a slower, more correct detokenizer on the output path rather than risk corrupt text. The exception is structured output, where the tokenizer's vocabulary meets a grammar and the constraint machinery can dominate the step.
3

FlashAttention tiling and online softmax

Never materialise the T×T matrix

Textbook attention computes scores S = QKᵀ, takes a row-wise softmax, then multiplies by V. For a sequence of length T that means writing and re-reading a T×T matrix in HBM — at T = 8K and FP16 that is 128 MB per attention head, per layer. FlashAttention avoids it by tiling: load a block of Q into SRAM, stream the K/V blocks past it, and maintain the softmax incrementally. The trick is the online softmax: for each new block you track the running maximum m and the running denominator l, rescaling the accumulator by e^{m_{\text{old}} - m_{\text{new}}} whenever the max moves. The result is mathematically identical to the full softmax, with HBM traffic that is O(T \cdot d) instead of O(T^2).

$$m_i \leftarrow \max(m_{i-1},\, \tilde m_i), \qquad l_i \leftarrow l_{i-1}\,e^{\,m_{i-1}-m_i} + \tilde l_i\,e^{\,\tilde m_i-m_i}$$

The animation sweeps the K/V tiles for one query tile and prints the live (m, l) pair and the resulting HBM traffic for the sequence length you choose. Notice that m only ever rises and l only ever needs rescaling when it does — the incremental update is exact, not an approximation.

Tile sweep is animated; sequence length is a slider. Press pause to inspect a tile.

Tiling and online-softmax formulation follow Dao et al., "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness" (NeurIPS 2022). HBM byte counts here are a per-head model that ignores L2 reuse, so treat the ratio, not the absolute figure, as the point.

4

FlashDecoding and split-K

Prefill has a full grid; decode at batch 1 has almost none

FlashAttention assumes there are enough query tiles to fill the GPU. Prefill has T of them, so it parallelises trivially. A decode step at batch 1 has exactly one query token and hundreds of KV blocks: the attention kernel is a handful of warps long and the GPU sits idle. FlashDecoding fixes this by splitting the K/V axis, running a partial attention over each split with the same online-softmax state, and combining the partials in a second reduction kernel. This is the same split-K idea as a GEMM with a skinny M dimension, applied to attention. It buys several times the decode throughput at low batch and changes nothing about the result. At high batch there are already enough query rows to fill the grid, so engines disable the split — another crossover that runs the opposite way to the CPU-launch one below.

💡 The shape of the whole page: prefill is a big parallel GEMM that likes tiling and graphs are unnecessary; decode at low batch is a latency-bound scatter of tiny kernels that needs split-K and CUDA graphs to avoid starving the GPU. A serving engine is two different machines wearing one interface.
5

Paged kernels and FlashInfer

Attention over a block table, not a contiguous tensor

PagedAttention makes the KV cache a set of non-contiguous 16-token blocks, which is what allows sharing and fragmentation control in the first place. The cost is that the attention and KV-write kernels can no longer assume a dense tensor: they gather blocks through a block table, and the memory access pattern is irregular. Writing those kernels by hand is where most of an engine's low-level effort goes, and it is why engines cluster around a few kernel libraries. FlashInfer is the clearest example — paged and ragged batched kernels with customisable attention variants, JIT-compiled for the shapes the server actually sees — and it is the backend underneath more than one engine. The practical consequence for capacity planning is that paged kernels trade a small amount of throughput for a large amount of memory flexibility, and the trade is almost always worth it on real traffic.

6

CPU launch overhead at small batch

At batch 1 the GPU waits for the CPU

A transformer layer is not one kernel; it is a dozen or more — RMSNorm, QKV projection, RoPE, attention, output projection, MLP gate/up/down, plus sampling. A 70B model over 80 layers issues on the order of a thousand kernel launches per decode step. Each launch costs a few microseconds of host time, and the GPU cannot start work it has not been handed. At large batch each kernel runs for milliseconds, so the host is far ahead and launch overhead vanishes into the noise. At batch 1 the kernels are tiny, the host becomes the bottleneck, and the measured step time is set by the CPU rather than the GPU. The batch slider in step 1 is exactly this curve: roughly 45% of a batch-1 step is CPU launch, falling to about 4% by batch 128. This is why every serious engine has a CUDA-graph path that exists purely to delete the host from the inner loop.

7

CUDA graphs and torch.compile

Record the launch stream once, replay it as one operation

A CUDA graph captures a sequence of kernels and their dependencies once, then replays the whole sequence with a single launch. The host disappears from the critical path: the GPU executes back-to-back and the inter-kernel gaps collapse. The constraints are real — graph capture needs fixed shapes and fixed memory addresses, so the engine pads and buckets sequence lengths and uses a static memory pool, and any control flow that depends on data must be turned into a branch that the graph already contains. In practice engines capture a graph per batch-size bucket and per sequence-length bucket, which is why "graph capture" is a distinct cold-start stage (Part 19) and why changing a shape can force a recapture. torch.compile attacks the same problem from the other side: it traces and fuses the Python-level model code into fewer, larger kernels, reducing both launch count and memory traffic, and the two compose — a compiled model is what you capture into a graph.

CPU launch stream above, GPU execution below; toggle the graph and watch the gaps close.

8

Sampling kernels

A few microseconds each, until they are not

Sampling looks free next to a matmul and stops being free once you stack processors: temperature, top-k, top-p, min-p, repetition penalty, log-probability extraction, and for structured output a token mask built from a grammar automaton. The mask is the expensive one — it can be tens to hundreds of microseconds per step depending on grammar complexity — and it is the reason constrained decoding has a measurable tax. The semantics of each processor are the training guide's subject; the link below goes there rather than re-deriving them. What belongs here is the cost model, because the sampler runs every single decode step and therefore competes directly with the kernels it feeds.

Toggle processors; bar height is added microseconds per decode step.

📌 Cross-link: what temperature, top-k, top-p, min-p and repetition penalty actually do to the distribution is derived interactively in the training guide's sampler; this page only counts their cost.
9

The vLLM V1 split and the engine landscape

EngineCore, AsyncLLM, and who supports what

vLLM's V1 architecture made the process split explicit: AsyncLLM owns the API surface, tokenization, detokenization and streaming, while EngineCore owns the model, the scheduler and the GPU loop, with the two communicating over ZMQ. The split exists to keep the GPU iteration tight — a slow client or a large JSON payload never blocks a step — and it is representative of where the field converged: the frontier is not the transformer, it is everything scheduled around it. TensorRT-LLM pushes the same ideas into NVIDIA's compiled stack with FP8/FP4 and in-flight batching; SGLang pairs a RadixAttention prefix tree with a strong structured-output front end; llama.cpp trades paged attention and disaggregation for portability across CPUs, Metal and consumer GPUs. The table is generated from the guide's engine dataset — select features to see who actually ships them.

✓

Cheat sheet

The engine in one table

PieceWhy it existsNumber to remember
FlashAttention tilingNever write the T×T score matrix to HBMHBM traffic O(T·d) vs O(T²)
Online softmaxExact softmax from streamed blocksRescale by e^(m_old − m_new) when the max rises
FlashDecoding split-KFill the GPU when decode has one query rowSplit the KV axis, reduce partials
Paged / FlashInfer kernelsAttention over non-contiguous blocks16-token blocks gathered via a block table
CPU launch overheadHost issues ~10³ kernels per step~45% of a batch-1 step → ~4% at batch 128
CUDA graphReplay the whole launch stream as one opFixed shapes, captured per bucket
Sampler costProcessors run every decode stepFew µs each; grammar masks 40–150 µs
EngineCore / AsyncLLMKeep the GPU loop off the API pathScheduler + model in a separate process
📚

Further reading

References

?

Check your understanding

0/5 answered