Inside the engine: kernels, CUDA graphs and the life of a request
A scheduler decides what to run; an engine decides how. This part follows a single request across the process boundary into the GPU worker, then opens the worker up: the tokenizer and detokenizer on the way in and out, the FlashAttention and paged kernels in the middle, the CPU that launches them, and the CUDA graph and compilation layers that exist because at small batch the launches cost more than the maths. It ends with the landscape of engines that implement all of this — and what each one actually supports.
The end-to-end path across process boundaries
One request, six handoffs, at least two processes
The training guide's serving page stops at the GPU: prefill then decode, batched continuously. A real engine has to move text across process boundaries before and after that. A gateway terminates HTTP, authenticates and enforces quotas; a tokenizer process converts text to ids; the scheduler process — vLLM V1 calls it EngineCore — owns the model and decides what runs each iteration; the GPU worker executes kernels; the detokenizer converts ids back; and a streaming layer flushes deltas to the client. The reason this matters at serving scale is that every handoff adds latency the GPU cannot hide, and the naive implementation serialises them. Production engines pipeline the stages and keep the model loop in a separate process precisely so a slow JSON parse or a blocking socket write cannot stall a decode iteration.
Click any stage in the strip. The batch slider at the bottom is the thread that ties the whole part together: the fraction of a step spent launching kernels on the CPU collapses as batch size grows, which is why the expensive machinery in the second half of this page only starts to pay at small batch.
Click a stage; drag the batch slider to move the CPU/GPU split.
The tokenizer, the detokenizer and the streaming path
The bytes you actually stream are not tokens
Tokenization is CPU work measured in tens of microseconds for a normal prompt — irrelevant next to a prefill — but it is on the critical path, and a pathological prompt (long, emoji-heavy, or byte-fallback-heavy) can cost tens of milliseconds. The output side is subtler. A token is a byte sequence, not necessarily valid UTF-8, so a naive detokenizer that emits each token as it is sampled can split a multi-byte character across two SSE events and produce mojibake. Correct streaming detokenizers buffer until a token boundary is a valid character boundary, apply stop-string matching across deltas, and only then flush. That buffering is why "streaming" in a well-built engine is not token-by-token: it is the smallest safe text delta the client can append.
FlashAttention tiling and online softmax
Never materialise the T×T matrix
Textbook attention computes scores S = QKᵀ, takes a row-wise softmax, then multiplies by V. For a sequence of length T that means writing and re-reading a T×T matrix in HBM — at T = 8K and FP16 that is 128 MB per attention head, per layer. FlashAttention avoids it by tiling: load a block of Q into SRAM, stream the K/V blocks past it, and maintain the softmax incrementally. The trick is the online softmax: for each new block you track the running maximum m and the running denominator l, rescaling the accumulator by e^{m_{\text{old}} - m_{\text{new}}} whenever the max moves. The result is mathematically identical to the full softmax, with HBM traffic that is O(T \cdot d) instead of O(T^2).
The animation sweeps the K/V tiles for one query tile and prints the live (m, l) pair and the resulting HBM traffic for the sequence length you choose. Notice that m only ever rises and l only ever needs rescaling when it does — the incremental update is exact, not an approximation.
Tile sweep is animated; sequence length is a slider. Press pause to inspect a tile.
Tiling and online-softmax formulation follow Dao et al., "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness" (NeurIPS 2022). HBM byte counts here are a per-head model that ignores L2 reuse, so treat the ratio, not the absolute figure, as the point.
FlashDecoding and split-K
Prefill has a full grid; decode at batch 1 has almost none
FlashAttention assumes there are enough query tiles to fill the GPU. Prefill has T of them, so it parallelises trivially. A decode step at batch 1 has exactly one query token and hundreds of KV blocks: the attention kernel is a handful of warps long and the GPU sits idle. FlashDecoding fixes this by splitting the K/V axis, running a partial attention over each split with the same online-softmax state, and combining the partials in a second reduction kernel. This is the same split-K idea as a GEMM with a skinny M dimension, applied to attention. It buys several times the decode throughput at low batch and changes nothing about the result. At high batch there are already enough query rows to fill the grid, so engines disable the split — another crossover that runs the opposite way to the CPU-launch one below.
Paged kernels and FlashInfer
Attention over a block table, not a contiguous tensor
PagedAttention makes the KV cache a set of non-contiguous 16-token blocks, which is what allows sharing and fragmentation control in the first place. The cost is that the attention and KV-write kernels can no longer assume a dense tensor: they gather blocks through a block table, and the memory access pattern is irregular. Writing those kernels by hand is where most of an engine's low-level effort goes, and it is why engines cluster around a few kernel libraries. FlashInfer is the clearest example — paged and ragged batched kernels with customisable attention variants, JIT-compiled for the shapes the server actually sees — and it is the backend underneath more than one engine. The practical consequence for capacity planning is that paged kernels trade a small amount of throughput for a large amount of memory flexibility, and the trade is almost always worth it on real traffic.
CPU launch overhead at small batch
At batch 1 the GPU waits for the CPU
A transformer layer is not one kernel; it is a dozen or more — RMSNorm, QKV projection, RoPE, attention, output projection, MLP gate/up/down, plus sampling. A 70B model over 80 layers issues on the order of a thousand kernel launches per decode step. Each launch costs a few microseconds of host time, and the GPU cannot start work it has not been handed. At large batch each kernel runs for milliseconds, so the host is far ahead and launch overhead vanishes into the noise. At batch 1 the kernels are tiny, the host becomes the bottleneck, and the measured step time is set by the CPU rather than the GPU. The batch slider in step 1 is exactly this curve: roughly 45% of a batch-1 step is CPU launch, falling to about 4% by batch 128. This is why every serious engine has a CUDA-graph path that exists purely to delete the host from the inner loop.
CUDA graphs and torch.compile
Record the launch stream once, replay it as one operation
A CUDA graph captures a sequence of kernels and their dependencies once, then replays the whole sequence with a single launch. The host disappears from the critical path: the GPU executes back-to-back and the inter-kernel gaps collapse. The constraints are real — graph capture needs fixed shapes and fixed memory addresses, so the engine pads and buckets sequence lengths and uses a static memory pool, and any control flow that depends on data must be turned into a branch that the graph already contains. In practice engines capture a graph per batch-size bucket and per sequence-length bucket, which is why "graph capture" is a distinct cold-start stage (Part 19) and why changing a shape can force a recapture. torch.compile attacks the same problem from the other side: it traces and fuses the Python-level model code into fewer, larger kernels, reducing both launch count and memory traffic, and the two compose — a compiled model is what you capture into a graph.
CPU launch stream above, GPU execution below; toggle the graph and watch the gaps close.
Sampling kernels
A few microseconds each, until they are not
Sampling looks free next to a matmul and stops being free once you stack processors: temperature, top-k, top-p, min-p, repetition penalty, log-probability extraction, and for structured output a token mask built from a grammar automaton. The mask is the expensive one — it can be tens to hundreds of microseconds per step depending on grammar complexity — and it is the reason constrained decoding has a measurable tax. The semantics of each processor are the training guide's subject; the link below goes there rather than re-deriving them. What belongs here is the cost model, because the sampler runs every single decode step and therefore competes directly with the kernels it feeds.
Toggle processors; bar height is added microseconds per decode step.
The vLLM V1 split and the engine landscape
EngineCore, AsyncLLM, and who supports what
vLLM's V1 architecture made the process split explicit: AsyncLLM owns the API surface, tokenization, detokenization and streaming, while EngineCore owns the model, the scheduler and the GPU loop, with the two communicating over ZMQ. The split exists to keep the GPU iteration tight — a slow client or a large JSON payload never blocks a step — and it is representative of where the field converged: the frontier is not the transformer, it is everything scheduled around it. TensorRT-LLM pushes the same ideas into NVIDIA's compiled stack with FP8/FP4 and in-flight batching; SGLang pairs a RadixAttention prefix tree with a strong structured-output front end; llama.cpp trades paged attention and disaggregation for portability across CPUs, Metal and consumer GPUs. The table is generated from the guide's engine dataset — select features to see who actually ships them.
Cheat sheet
The engine in one table
| Piece | Why it exists | Number to remember |
|---|---|---|
| FlashAttention tiling | Never write the T×T score matrix to HBM | HBM traffic O(T·d) vs O(T²) |
| Online softmax | Exact softmax from streamed blocks | Rescale by e^(m_old − m_new) when the max rises |
| FlashDecoding split-K | Fill the GPU when decode has one query row | Split the KV axis, reduce partials |
| Paged / FlashInfer kernels | Attention over non-contiguous blocks | 16-token blocks gathered via a block table |
| CPU launch overhead | Host issues ~10³ kernels per step | ~45% of a batch-1 step → ~4% at batch 128 |
| CUDA graph | Replay the whole launch stream as one op | Fixed shapes, captured per bucket |
| Sampler cost | Processors run every decode step | Few µs each; grammar masks 40–150 µs |
| EngineCore / AsyncLLM | Keep the GPU loop off the API path | Scheduler + model in a separate process |
Further reading
References
- Dao et al., "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness" (NeurIPS 2022).
- Dao et al., "Flash-Decoding for long-context inference" (2023) — split-K over the KV axis.
- Ye et al., "FlashInfer: Efficient and Customizable Attention Engine for LLM Inference" (2025).
- Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention" (SOSP 2023).
- vLLM project, V1 engine architecture documentation — AsyncLLM and EngineCore.
- Dong et al., "XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models" (MLSys 2025).