Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

The life of a request

Click any stage to jump to the part that covers it. The request path runs left to right across lane A; lane B is what a single worker is made of; lane C is the substrate underneath all of it.

A · the request path Client / API gateway Parts 17, 19 429s, token buckets KV-aware load balancer Parts 8, 19 prefix affinity Queue + scheduler Parts 6, 7 SLOs, fairness Prefill worker Parts 1, 6, 12 compute-bound Decode worker Parts 1, 6, 12 memory-bound KV transfer over RDMA · Part 13 next step / preemption · Part 7 B · what a worker is made of Paged KV cache Part 5 16-token blocks Prefix cache / radix Part 8 shared prefixes Attention & sampling Part 9 FlashAttention tiling Quantized weights Part 10 FP8 / MXFP4 Draft model Part 11 EAGLE-3 TP · PP · EP shards Part 12 NVLink domain C · the substrate GPU + HBM bandwidth Part 2 ~1–2 FLOPs/byte decode MoE experts & all-to-all Part 14 top 4% take 38% of tokens K8s autoscaling / cold start Part 19 40–90 s to first token SLOs, goodput, $/M tokens Parts 3, 20 cost per task, not per token Workload shapes Long context Part 15 quadratic prefill, linear KV Structured output Part 16 150K-wide token mask Agents & multimodal Part 18 prefill:decode inversion Reasoning Parts 18, 20 thinking tokens blow up outputs

All 20 parts

Part 1 of 20

The decode loop and the memory wall

Why generating one token costs a full sweep of the weights through HBM, and why that single fact shapes everything that follows.

Part 2 of 20

The machine: HBM, interconnect and the accelerator zoo

Bandwidth as the headline number, the all-reduce tax, NVL72 as one big GPU, SRAM architectures, and energy per token.

Part 3 of 20

TTFT, TPOT and goodput

The four numbers that describe a serving system, why percentiles beat means, and how to write an SLO you can defend.

Part 4 of 20

The KV cache: sizing, GQA, MQA and MLA

The bytes-per-token formula run against real model configs, and the lever it becomes for concurrency.

Part 5 of 20

PagedAttention: block tables, fragmentation and preemption

The three kinds of waste in contiguous reservation, blocks and block tables, copy-on-write, and recompute versus swap.

Part 6 of 20

Continuous batching and chunked prefill

Why iteration-level batching wins, what batch size does to step time, and the ITL spike chunked prefill trades away.

Part 7 of 20

The scheduler: queues, fairness and admission control

What the scheduler decides every step, head-of-line blocking, length prediction, starvation, and when to shed load.

Part 8 of 20

Prefix caching, radix trees and the KV hierarchy

Reusing the same prefix across requests, RadixAttention, cache-aware routing, and the tier economics of fetch versus recompute.

Part 9 of 20

Inside the engine: kernels, CUDA graphs and the life of a request

The end-to-end request path, FlashAttention tiling, paged kernels, launch overhead, and the engine landscape.

Part 10 of 20

Quantization and numerics

Formats from FP16 to MXFP4, activation outliers, why the win is bandwidth not FLOPs, and how to measure the damage.

Part 11 of 20

Speculative decoding

Verify many, generate one; acceptance rates, draft models, tree attention, and why engines disable speculation under load.

Part 12 of 20

Parallelism for inference: TP, PP, EP, DP-attention, CP

Inference parallelism is not training parallelism, why TP stays inside NVLink, and how to search the sharding space.

Part 13 of 20

Prefill/decode disaggregation

One pool, two jobs; sizing the pools, moving KV over RDMA, when disaggregation loses, and what DeepSeek actually runs.

Part 14 of 20

Serving mixture-of-experts at scale

Active versus total parameters, all-to-all dispatch, hot experts, EPLB, wide expert parallelism, and why rack-scale exists.

Part 15 of 20

Long-context serving

Quadratic prefill, linear KV, ring attention, sparse attention and eviction, and the honest cost per request.

Part 16 of 20

Structured output and constrained decoding

Regex to FSM to per-state masks, why JSON needs a pushdown automaton, jump-forward decoding, and the constraint tax.

Part 17 of 20

LoRA, quotas and noisy neighbours

One base and many adapters, batched LoRA kernels, adapter paging, token buckets, and fair queueing across tenants.

Part 18 of 20

Reasoning, agents and multimodal

Workload shape as a first-class serving parameter: thinking-token budgets, agent turns, image tokens, and realtime voice.

Part 19 of 20

Kubernetes, autoscaling, cold starts and routing

Gateways, routers and autoscalers; why GPU utilisation is a lie; the cold-start breakdown; and honest scale-to-zero.

Part 20 of 20

Benchmarking, cost and capacity planning

Open versus closed loop, batch non-invariance, dashboard signals, incident signatures, and RPS to GPUs to dollars.

Reference

Glossary & numbers to know

Every term in the series linked back to where it is introduced, plus the accelerator, model, quantization, engine and pricing tables.

Already know some of this? Start here

If you...Start at
Read Part 11 of the training guide and want the real thingPart 1 — The decode loop and the memory wall
Just need to size a GPU for a modelPart 2 — The machine and Part 4 — The KV cache
p99 TTFT is bad but throughput looks finePart 3 — TTFT, TPOT and goodput, then Parts 6–7
Are paying too much per tokenPart 20 — Benchmarking, cost and capacity planning
Serve agents and the KV cache thrashesPart 8 — Prefix caching, then Part 18 — Workloads
Want to know what DeepSeek actually runsPart 13 — Prefill/decode disaggregation
Are deploying on Kubernetes tomorrowPart 19 — Kubernetes, autoscaling, cold starts and routing
Need to guarantee valid JSON or a regexPart 16 — Structured output and constrained decoding
Start from the beginning, or jump straight into whichever part you need. Start with Part 1 →