LLM Serving, Interactively
Part 11 of the training guide established the KV cache, prefill versus decode, and continuous batching in twenty minutes. This series rebuilds all of it as a quantitative model you plug your own hardware into, then follows it out to what hyperscalers actually run: disaggregated prefill and decode pools, wide expert parallelism, rack-scale NVLink, and a cost per task you can defend.
The life of a request
Click any stage to jump to the part that covers it. The request path runs left to right across lane A; lane B is what a single worker is made of; lane C is the substrate underneath all of it.
All 20 parts
The decode loop and the memory wall
Why generating one token costs a full sweep of the weights through HBM, and why that single fact shapes everything that follows.
Part 2 of 20The machine: HBM, interconnect and the accelerator zoo
Bandwidth as the headline number, the all-reduce tax, NVL72 as one big GPU, SRAM architectures, and energy per token.
Part 3 of 20TTFT, TPOT and goodput
The four numbers that describe a serving system, why percentiles beat means, and how to write an SLO you can defend.
Part 4 of 20The KV cache: sizing, GQA, MQA and MLA
The bytes-per-token formula run against real model configs, and the lever it becomes for concurrency.
Part 5 of 20PagedAttention: block tables, fragmentation and preemption
The three kinds of waste in contiguous reservation, blocks and block tables, copy-on-write, and recompute versus swap.
Part 6 of 20Continuous batching and chunked prefill
Why iteration-level batching wins, what batch size does to step time, and the ITL spike chunked prefill trades away.
Part 7 of 20The scheduler: queues, fairness and admission control
What the scheduler decides every step, head-of-line blocking, length prediction, starvation, and when to shed load.
Part 8 of 20Prefix caching, radix trees and the KV hierarchy
Reusing the same prefix across requests, RadixAttention, cache-aware routing, and the tier economics of fetch versus recompute.
Part 9 of 20Inside the engine: kernels, CUDA graphs and the life of a request
The end-to-end request path, FlashAttention tiling, paged kernels, launch overhead, and the engine landscape.
Part 10 of 20Quantization and numerics
Formats from FP16 to MXFP4, activation outliers, why the win is bandwidth not FLOPs, and how to measure the damage.
Part 11 of 20Speculative decoding
Verify many, generate one; acceptance rates, draft models, tree attention, and why engines disable speculation under load.
Part 12 of 20Parallelism for inference: TP, PP, EP, DP-attention, CP
Inference parallelism is not training parallelism, why TP stays inside NVLink, and how to search the sharding space.
Part 13 of 20Prefill/decode disaggregation
One pool, two jobs; sizing the pools, moving KV over RDMA, when disaggregation loses, and what DeepSeek actually runs.
Part 14 of 20Serving mixture-of-experts at scale
Active versus total parameters, all-to-all dispatch, hot experts, EPLB, wide expert parallelism, and why rack-scale exists.
Part 15 of 20Long-context serving
Quadratic prefill, linear KV, ring attention, sparse attention and eviction, and the honest cost per request.
Part 16 of 20Structured output and constrained decoding
Regex to FSM to per-state masks, why JSON needs a pushdown automaton, jump-forward decoding, and the constraint tax.
Part 17 of 20LoRA, quotas and noisy neighbours
One base and many adapters, batched LoRA kernels, adapter paging, token buckets, and fair queueing across tenants.
Part 18 of 20Reasoning, agents and multimodal
Workload shape as a first-class serving parameter: thinking-token budgets, agent turns, image tokens, and realtime voice.
Part 19 of 20Kubernetes, autoscaling, cold starts and routing
Gateways, routers and autoscalers; why GPU utilisation is a lie; the cold-start breakdown; and honest scale-to-zero.
Part 20 of 20Benchmarking, cost and capacity planning
Open versus closed loop, batch non-invariance, dashboard signals, incident signatures, and RPS to GPUs to dollars.
ReferenceGlossary & numbers to know
Every term in the series linked back to where it is introduced, plus the accelerator, model, quantization, engine and pricing tables.
Already know some of this? Start here
| If you... | Start at |
|---|---|
| Read Part 11 of the training guide and want the real thing | Part 1 — The decode loop and the memory wall |
| Just need to size a GPU for a model | Part 2 — The machine and Part 4 — The KV cache |
| p99 TTFT is bad but throughput looks fine | Part 3 — TTFT, TPOT and goodput, then Parts 6–7 |
| Are paying too much per token | Part 20 — Benchmarking, cost and capacity planning |
| Serve agents and the KV cache thrashes | Part 8 — Prefix caching, then Part 18 — Workloads |
| Want to know what DeepSeek actually runs | Part 13 — Prefill/decode disaggregation |
| Are deploying on Kubernetes tomorrow | Part 19 — Kubernetes, autoscaling, cold starts and routing |
| Need to guarantee valid JSON or a regex | Part 16 — Structured output and constrained decoding |