Glossary & numbers to know
This is the serving guide's back matter: one alphabetical index that links every term to the part that introduces it, the sourced tables the rest of the series cites, and a formula card for the back-of-envelope arithmetic. It is the serving-side counterpart to the training guide's glossary, which defines the modelling ideas; here the emphasis is the number, the failure mode and the deployment shape.
Accelerators: the bandwidth table
Capacity, bandwidth, FLOPs, fabric, power
A serving system is bounded below by the memory bandwidth of the chip it runs on. HBM capacity decides how many KV blocks fit and therefore how many sequences can be resident; bandwidth decides how fast a decode step can sweep the weights and the cache; peak FLOPs set the prefill ceiling; the interconnect decides whether you may shard at all. Read the table by asking which column is smallest for your workload. A single-stream decode job is almost always bandwidth-bound, so bwTBs is the number that predicts tokens per second; a long-prompt prefill job is compute-bound, so bf16TFLOPs and fp8TFLOPs matter more; a multi-node tensor-parallel deployment lives or dies on the fabric column.
The NVIDIA headline FLOPs in the upstream source include 2:4 sparsity, so the dense values rendered here are half the press-release number. AMD and Google publish dense figures directly. A dash means the format is not natively supported by that part, which is why FP4 is a Blackwell story and FP8 is a Hopper-and-later story.
| Accelerator | HBM | BW | BF16 | FP8 | FP4 | Interconnect | TDP |
|---|
Models and their KV configurations
The two numbers that drive a capacity plan
Model parameters set the weight bytes a decode step must read; KV bytes per token set how many sequences fit beside those weights. Both are computed from public configs, not estimated. For an MHA or GQA model the per-token cache is 2 × layers × n_kv_heads × d_head × 2 bytes — the leading 2 for K and V, the trailing 2 for FP16. MLA breaks the formula because it caches a single low-rank latent plus a decoupled RoPE key, and SWA-hybrid models bound the cache at the sliding-window length instead of letting it grow linearly, so read their rows as ceilings rather than rates. The KV bytes column is the one the concurrency calculators in the series use; multiply by context length and concurrency to see whether a model fits.
| Model | Layers | Q heads | KV heads | d_head | Scheme | KV per token (FP16) | Params total / active | Context |
|---|
Back-of-envelope formulas
Six expressions that explain most serving behaviour
Every interactive demo in the series is a version of one of these. Keep them in a notebook; they are accurate to a factor you can reason about, which is the point. Bytes per element is written β and equals 2 for BF16/FP16, 1 for FP8/INT8 and 0.5 for INT4.
KV bytes
L layers, Hkv KV heads, head width dhead, T cached tokens, batch b. For MLA replace 2 Hkv dhead with the cached latent width (576 for DeepSeek-V3).
Step time
The max of the memory-bound and compute-bound terms, plus launch and synchronization overhead. Decode at small batch is memory-bound; prefill is compute-bound.
Expected speculative tokens
Per draft round with acceptance rate α and γ draft tokens: 1 + α + α2 + … + αγ. At α=0.8 and γ=4 the round yields about 3.36 tokens.
Pipeline bubble fraction
p pipeline stages and m microbatches. The fraction of stage-time spent idle while the pipe fills and drains; it vanishes only as m grows large.
All-reduce bytes
Ring all-reduce moving message size S across N ranks. The factor 2 is the reduce-scatter plus all-gather; it is the per-step tax every TP group pays.
Goodput
Requests per second that satisfy both latency limits. Raw throughput counts work you cannot bill for; goodput counts only the work that met its contract.
Glossary
Every term, linked to the part that introduces it
| Term | What it means | Introduced in |
|---|---|---|
| A | ||
| Acceptance rate | Expected fraction of speculative draft tokens the target model accepts in a verification pass; the single number that decides whether speculation pays. | Part 11 |
| Accelerator zoo | The GPU/TPU/Trainium landscape compared by HBM, bandwidth, FLOPs and interconnect rather than by brand. | Part 2 |
| Activation outlier | A few input channels whose values dwarf the rest, setting a coarse step for every ordinary channel under per-tensor scaling. | Part 10 |
| Active parameters | Parameters actually touched per token in a mixture-of-experts model (37B of 671B for DeepSeek-V3), as distinct from the total resident. | Part 14 |
| Admission control | Deciding whether to accept a request into the running set or shed it, based on KV headroom, capacity and the SLO. | Part 7 |
| Aging | Increasing a waiting request's effective priority over time so that long jobs cannot starve behind a stream of short ones. | Part 7 |
| All-reduce | Summing a tensor across ranks and broadcasting the result; the synchronization every tensor-parallel step pays. | Part 12 |
| All-to-all | Every rank exchanges a different slice with every other rank; the dispatch/combine primitive of expert parallelism. | Part 14 |
| Arithmetic intensity | FLOPs performed per byte moved; about 1–2 for a decode step and far higher for prefill, which is why the two live in different regimes of the roofline. | Part 1 |
| Attention sink | The first few tokens that absorb a disproportionate share of attention and must be kept for stability when the cache is trimmed. | Part 15 |
| B | ||
| Backpressure | Returning a 429 or slowing the client instead of letting an unbounded queue convert overload into an SLO breach for everyone. | Part 7 |
| Batch invariance | Whether a request's output is bit-identical regardless of the other requests it was batched with; it is not, because floating-point addition is not associative. | Part 20 |
| Batch size (max_num_seqs) | How many sequences the engine runs per iteration; the knob that moves a step from memory-bound toward compute-bound. | Part 6 |
| Block manager | The allocator mapping a sequence's logical token positions to physical KV blocks, and the component preemption calls into. | Part 5 |
| Block size | Tokens per KV page, conventionally 16; the granularity at which internal fragmentation is paid. | Part 5 |
| Block table | The per-sequence array that translates logical block indices to physical block indices, so paged KV needs no contiguity. | Part 5 |
| Bottleneck (memory vs compute) | Which resource limits a step; decode is memory-bandwidth-bound, prefill is compute-bound, and batch size moves the crossing point. | Part 1 |
| C | ||
| Cache hit rate | Fraction of prompt tokens whose KV was already resident; the statistic that turns prefix caching from an optimisation into a throughput multiplier. | Part 8 |
| Cache-aware routing | Sending a request to the replica that already holds its prefix instead of the least-loaded one. | Part 8 |
| Calibration | Running representative data through the model to measure activation ranges before quantizing; the ranges are data-dependent, so this is not really post-training. | Part 10 |
| Capacity planning | Converting an RPS and latency target into GPUs and dollars, including the queueing headroom the mean hides. | Part 20 |
| Chunked prefill | Splitting a long prompt's prefill across iterations so decode requests are not stalled for the whole prompt's compute time. | Part 6 |
| Cold start | The 40–90 s from pod schedule to first token: image pull, process init, weight load, graph capture and warmup. | Part 19 |
| Compute-bound | Step time set by FLOPs rather than bytes; the high-batch and long-prefill regime. | Part 1 |
| Connector (KV) | The transport that ships a finished prefill's KV blocks to a decode worker, often over RDMA or NIXL. | Part 13 |
| Context length | Tokens the model can attend to; KV memory and decode cost grow linearly with it while prefill cost grows quadratically. | Part 4 |
| Context parallelism (CP) | Splitting the sequence dimension across devices so a long context's attention and KV fit, with activations passed around the ring. | Part 12 |
| Continuous batching | Admitting and retiring sequences every iteration rather than per fixed batch, so a slow request never idles a whole batch. | Part 6 |
| Copy-on-write | Two sequences share a physical KV block until one writes, at which point the block is copied; how shared prefixes are stored once. | Part 5 |
| Cost per task | Dollars per completed task rather than per token; the number a budget actually sees once retries and thinking tokens are counted. | Part 20 |
| CUDA graph | A captured sequence of kernel launches replayed as a single operation to remove per-launch CPU overhead from a decode step. | Part 9 |
| D | ||
| Data parallelism (DP) | Replicating the model and splitting requests across copies; the default until weights no longer fit one device. | Part 12 |
| Decode | Generating one token at a time, memory-bandwidth-bound because every step sweeps the weights through HBM. | Part 1 |
| Decode worker | The pool in a disaggregated deployment that holds KV and generates tokens after prefill has produced it. | Part 13 |
| DeepEP | DeepSeek's open-source MoE communication library implementing grouped all-to-all dispatch and combine. | Part 14 |
| Defragmentation | Reclaiming scattered free KV blocks, or moving pages so there is no need to; paging turns this from a stop-the-world op into an allocation detail. | Part 5 |
| Disaggregation (PD) | Running prefill and decode in separate pools, each sized for its own bottleneck, with KV transferred between them. | Part 13 |
| Distributed KV cache | Partitioning a long context's KV across devices, for example with ring attention, so no single GPU holds all of it. | Part 15 |
| DP-attention | Combining data parallelism with attention sharding so KV is not replicated across replicas, trading communication for capacity. | Part 12 |
| Draft model | The small model that proposes speculative tokens for the target model to verify; its agreement rate sets the speedup. | Part 11 |
| Draft tree | Several candidate continuations explored in parallel and verified together, as in EAGLE and Medusa. | Part 11 |
| DuoAttention | Keeping full KV for a small set of retrieval heads and constant-length KV for streaming heads; up to 2.18× decode speedup for MHA. | Part 15 |
| E | ||
| EAGLE | Speculative decoding that drafts in the target model's feature space; EAGLE-2 reports 3.05–4.26× and EAGLE-3 up to 6.5×. | Part 11 |
| End-to-end latency | Arrival to last token: queueing plus prefill plus decode, and the number a streaming UI hides behind TTFT. | Part 3 |
| Energy per token | Joules per generated token, roughly power × step time divided by tokens per step; the metric that makes rack power a serving variable. | Part 2 |
| Engine | The serving software — vLLM, SGLang, TensorRT-LLM, llama.cpp — that schedules requests and runs the kernels. | Part 9 |
| EPLB | Expert-parallel load balancer: replicates hot experts onto less-loaded ranks to even out MoE traffic; roughly a 20% recovery is cited for DeepSeek-V3-class deployments. | Part 14 |
| Eviction (KV) | Dropping low-value KV entries to cap cache size, whether by attention score, recency or query relevance. | Part 15 |
| Expert | One of many parallel MLPs that a router selects per token; the unit sharded by expert parallelism. | Part 14 |
| Expert parallelism (EP) | Sharding experts across devices, one or more per rank, and paying an all-to-all to dispatch tokens and combine results. | Part 14 |
| Expert router | The gate that scores a token and picks the top-k experts to handle it; its load distribution is what makes hot experts a problem. | Part 14 |
| F | ||
| Fair queueing | Scheduling that gives each tenant a bounded share under contention, so one workload cannot consume the fleet. | Part 17 |
| FCFS | First-come-first-served ordering; simple and starvation-free, but one long request blocks everything behind it. | Part 7 |
| First-token latency | Another name for TTFT, usefully split into queueing time and prefill time when diagnosing a regression. | Part 3 |
| FlashAttention | IO-aware attention that tiles Q, K and V through SRAM instead of materializing the full attention matrix in HBM. | Part 9 |
| FP4 | 4-bit floating point, native on Blackwell at 0.5 bytes per parameter; the format that makes B200's FP4 peak a headline. | Part 10 |
| FP8 (E4M3 / E5M2) | 8-bit floating point, native on Hopper and Blackwell; E4M3 with per-block scaling is DeepSeek-V3's serving format. | Part 10 |
| Fragmentation | Wasted KV memory from rounding a sequence up to a block boundary or from gaps between reserved regions. | Part 5 |
| FSM (finite state machine) | The automaton a regular expression compiles to; its current state indexes the set of legal next tokens. | Part 16 |
| G | ||
| Gateway | The front door that authenticates, rate-limits, observes and routes requests before any GPU sees them. | Part 19 |
| Goodput | Requests per second that meet their latency SLOs; raw throughput counts work you cannot bill for. | Part 3 |
| GPU utilization | Fraction of time SMs are busy; routinely near 100% while goodput is poor, which is why it is a lie as a serving signal. | Part 19 |
| Graph capture | Recording a decode iteration's kernel launches into a CUDA graph so replay costs one launch instead of dozens. | Part 9 |
| Graph reuse | Caching captured CUDA graphs across process starts to remove graph-capture time from a cold start. | Part 19 |
| GQA | Grouped-query attention: several query heads share one KV head, cutting KV by 8× at Llama-3-70B geometry with little quality loss. | Part 4 |
| H | ||
| H2O | KV eviction that keeps a small set of heavy-hitter tokens plus recent ones; up to 29× throughput over naive baselines. | Part 15 |
| HBM | High-bandwidth memory: its capacity caps concurrency, its bandwidth caps decode speed, and both are read from the accelerator table. | Part 2 |
| Head dimension | Width of one attention head's vectors; it multiplies directly into KV bytes per token. | Part 4 |
| Head-of-line blocking | One long request delaying all the shorter ones queued behind it, the failure FCFS invites. | Part 7 |
| Heavy hitter | A token that receives disproportionate attention and is therefore worth keeping when the cache must shrink. | Part 15 |
| Hidden state | The residual-stream vector each layer reads and writes; the tensor whose size the weights matrix multiplies per token. | Part 1 |
| Hot expert | An expert that receives far more tokens than average; the top few percent of experts absorb a large share of traffic, which EPLB targets. | Part 14 |
| I | ||
| ICI | The TPU interconnect; Google's analogue of NVLink and the fabric that decides whether a TPU shard communicates cheaply. | Part 2 |
| Image tokens | Vision input expanded into context tokens; large, prefill-heavy and a reason multimodal workloads invert the prefill/decode ratio. | Part 18 |
| In-flight batching | TensorRT-LLM's name for continuous batching: sequences enter and leave the running set mid-flight. | Part 9 |
| Interconnect | The link (NVLink, NVSwitch, ICI, NeuronLink) that determines whether tensor and expert parallelism pay for themselves. | Part 2 |
| Inter-token latency (ITL) | The gap between consecutive output tokens; the inverse of TPOT and the number a chat user actually perceives as speed. | Part 3 |
| J | ||
| JSON schema | The grammar that constrains generation to a valid structure; the commonest reason structured decoding is deployed. | Part 16 |
| Jump-forward decoding | Emitting a run of forced tokens at once when the grammar admits only one continuation, rather than one step at a time. | Part 16 |
| K | ||
| Kernel | A compiled GPU routine; the unit whose launch overhead CUDA graphs remove and whose shape determines achieved efficiency. | Part 9 |
| Kubernetes (K8s) | The orchestrator that schedules serving pods, runs probes and drives autoscaling, and therefore sets the cold-start floor. | Part 19 |
| KV block | A fixed-size chunk of KV cache, the unit PagedAttention allocates and the block table maps. | Part 5 |
| KV cache | Cached keys and values so decode reuses attention state instead of recomputing it; the tensor that caps concurrency. | Part 4 |
| KV cache hit | A prefix whose KV was already computed and can be reused; a hit converts prefill work into a memory read. | Part 8 |
| KV compression | Storing or summarizing fewer, cheaper KV entries, whether by quantization, eviction or low-rank projection. | Part 15 |
| KV offload | Moving KV to a lower tier (DRAM or NVMe) and fetching it back when needed, trading latency for resident capacity. | Part 8 |
| KV quantization | Storing KV in FP8, INT8 or INT4 to fit more sequences; every error is re-read on every decode step, so evaluate, do not assume. | Part 10 |
| KV transfer | Shipping KV blocks over RDMA or NIXL from a prefill worker to a decode worker; the cost that bounds disaggregation. | Part 13 |
| KV-aware routing | Load balancing by which replica holds the request's prefix, not merely by which has spare capacity. | Part 8 |
| L | ||
| Latency SLO | The maximum TTFT or TPOT a request may have and still count as served; goodput is measured against it. | Part 3 |
| Length prediction | Guessing an output's length to schedule and admit better; wrong guesses either strand capacity or miss deadlines. | Part 7 |
| LLGuidance | A constrained-decoding library that compiles grammars into token masks for fast structured generation. | Part 16 |
| Load balancer | Distributes requests across replicas; a prefix-aware balancer turns a cache hit into a cheaper request. | Part 19 |
| Load shedding | Rejecting work deliberately to protect the requests already admitted; the honest alternative to an unbounded queue. | Part 7 |
| Logit | The raw pre-softmax score for each vocabulary entry; sampling turns the logit vector into one token. | Part 1 |
| Long-context | Serving very long prompts: quadratic prefill work, linear KV, and the eviction and routing machinery that makes it affordable. | Part 15 |
| LoRA | A low-rank adapter over frozen base weights; serving many of them over one base is the multi-tenant pattern. | Part 17 |
| M | ||
| Mask (logit) | Setting disallowed tokens to −∞ before sampling so the grammar is never violated. | Part 16 |
| MBU | Memory-bandwidth utilization: achieved bytes per second over peak; the utilisation that actually predicts decode speed. | Part 2 |
| Medusa | Adding decoding heads to a backbone so several future tokens are predicted and verified in parallel; 2.3–3.6× reported. | Part 11 |
| Memory wall | The widening gap between compute growth and bandwidth growth that makes decode memory-bound rather than compute-bound. | Part 1 |
| MFU | Model-FLOPs utilization: achieved FLOPs per second over the hardware peak; the prefill-side analogue of MBU. | Part 2 |
| MLA | Multi-head latent attention: DeepSeek's low-rank KV that caches one compressed latent plus a RoPE key, about 57× smaller than MHA at the same geometry. | Part 4 |
| Model Streamer | NVIDIA Run:ai's parallel weight loader; its benchmark drops a 15 GB model's load from 43.7 s to 7.5 s as concurrency rises. | Part 19 |
| MoE | Mixture of experts: many parameters resident but only a few active per token, which decouples capacity from FLOPs. | Part 14 |
| Mooncake | Moonshot's KVCache-centric disaggregated store; its vLLM integration reports 3.8× throughput and 46× lower P50 TTFT on agentic traces. | Part 8 |
| MQA | Multi-query attention: a single KV head, up to 64× less cache than MHA at the cost of quality headroom. | Part 4 |
| Multi-LoRA | Serving many adapters over one base model in a single batch, with adapters paged in and out like KV blocks. | Part 17 |
| Multi-tenancy | Sharing serving capacity across independent workloads or customers, with fairness and isolation rather than one global queue. | Part 17 |
| MXFP4 | Microscaling 4-bit with a shared block exponent; the weight format named by gpt-oss-120b and accelerated natively on Blackwell. | Part 10 |
| N | ||
| N-gram / prompt lookup | Speculating the continuation by copying matching n-grams from the prompt; no draft model, and the gain is entirely workload-dependent. | Part 11 |
| NIXL | NVIDIA's inference transfer library for moving KV between disaggregated workers over RDMA and other fabrics. | Part 13 |
| Noisy neighbour | A tenant whose load degrades the latency of others sharing the same replicas; the reason fair queueing is a serving feature. | Part 17 |
| Node (prefill / decode) | The unit of disaggregated capacity; DeepSeek runs 32-GPU prefill units against 144-GPU decode units in its published split. | Part 13 |
| NVFP4 | Blackwell's 4-bit floating-point format adding a per-block FP8 scale on top of E2M1 values. | Part 10 |
| O | ||
| OOM | Out-of-memory from KV growth; paging, preemption and admission control exist to prevent it from taking down a worker. | Part 5 |
| Open vs closed loop | Whether the client backs off when latency rises (closed) or pushes regardless (open); benchmarks must state which they use. | Part 20 |
| Outlines | The early FSM-indexing structured-generation library which introduced jump-forward decoding and near-zero masking overhead. | Part 16 |
| Overload | The regime where arrivals exceed capacity; the regime admission control, load shedding and backpressure exist for. | Part 7 |
| P | ||
| Page (KV) | A physical block of KV cache, the unit PagedAttention maps from a sequence's logical positions. | Part 5 |
| Paged KV cache | Non-contiguous, block-paged KV storage that removes the need to reserve a sequence's maximum length up front. | Part 5 |
| PagedAttention | The vLLM mechanism that pages KV and translates logical to physical blocks with a block table. | Part 5 |
| Padding | Wasted batch slots from forcing variable-length sequences into a rectangle; packing and paging avoid it. | Part 6 |
| Parallelism | Splitting a model's compute across devices: tensor, pipeline, expert, data (attention) and context parallelism. | Part 12 |
| Percentile (P50 / P90 / P99) | The latency below which that fraction of requests fall; the mean hides the tail that actually breaks SLOs. | Part 3 |
| Pipeline bubble | Idle time while pipeline stages fill and drain, proportional to (p−1)/(m+p−1) for p stages and m microbatches. | Part 12 |
| Pipeline parallelism (PP) | Splitting model layers across stages and streaming microbatches through them; needs less bandwidth than TP but adds latency and bubbles. | Part 12 |
| Prefill | Processing the whole prompt in parallel; compute-bound, and the source of TTFT. | Part 1 |
| Prefill-decode ratio | The mix of prompt and output tokens in a workload; agentic and multimodal traffic invert the ratio that chat assumes. | Part 18 |
| Prefill worker | The pool in a disaggregated deployment that only computes prompt KV and hands it to decoders. | Part 13 |
| Preemption | Evicting a running sequence's KV to make room, then recomputing or swapping it back when capacity returns. | Part 5 |
| Prefix affinity | The routing preference for the replica that already holds a shared prefix, so a cache hit is not thrown away. | Part 8 |
| Prefix caching | Reusing KV for shared prompt prefixes across requests; a cache hit turns prefill time into a memory read. | Part 8 |
| Priority scheduling | Ordering the queue by request class rather than arrival, which needs aging or quotas to stay starvation-free. | Part 7 |
| Pushdown automaton | A stack machine needed for nested grammars such as balanced JSON, since a finite-state machine cannot count depth. | Part 16 |
| PyramidKV | Layer-wise pyramidal KV budgets, more retention in low layers and less in high ones; near-full quality at roughly 12% of KV. | Part 15 |
| Q | ||
| Quantization | Storing weights, activations or KV at lower precision to save bytes; in serving the win is bandwidth, not FLOPs. | Part 10 |
| Query-aware eviction | Choosing KV entries to keep based on the current query rather than only recency or global attention mass. | Part 15 |
| Quest | Query-aware top-K KV page selection using per-page key ranges; up to 7.03× reported latency reduction. | Part 15 |
| R | ||
| RadixAttention | SGLang's radix-tree prefix cache; reported around 6× higher throughput on structured workloads. | Part 8 |
| Radix tree | The prefix tree that indexes cached KV by shared token sequences and makes longest-prefix matching cheap. | Part 8 |
| Recompute | Rebuilding a preempted sequence's KV from its tokens instead of moving it; cheaper than a fetch below roughly 1,900 tokens. | Part 5 |
| Rejection sampling | The rule that accepts a draft token only when the target model's sample agrees, keeping the output distribution exact. | Part 11 |
| Request routing | Choosing a replica for each request; the choice changes cache hit rate, queue depth and therefore latency. | Part 19 |
| Ring attention | Distributing a long sequence's attention across devices in a ring so each holds part of the KV and passes activations on. | Part 15 |
| Roofline | The model relating attainable FLOPs to arithmetic intensity, peak compute and peak bandwidth; it explains the memory wall. | Part 1 |
| RPS | Requests per second; the input to capacity planning once each request's token profile is known. | Part 20 |
| S | ||
| Sampling | Turning logits into a token with greedy, temperature or top-p rules; it is the last step of every decode iteration. | Part 1 |
| Scale-to-zero | Removing the last replica and paying a cold start to return; honest only when the traffic is genuinely bursty. | Part 19 |
| Scheduler | The component deciding what to admit, preempt and run each iteration, under KV and capacity limits. | Part 7 |
| Semantic caching | Reusing an answer when a new query is semantically equivalent to an old one; a workload-level cache above the KV cache. | Part 18 |
| Serving SLO | The latency and availability contract an API makes, expressed on TTFT, TPOT and error rate rather than on averages. | Part 3 |
| SGMV | A batched LoRA kernel that applies many adapters in one pass by treating the batch as a segmented matrix-vector product. | Part 17 |
| Sharding | Splitting weights, experts or KV across devices; every sharding choice trades memory for communication. | Part 12 |
| SJF | Shortest-job-first ordering; improves mean latency but starves long jobs unless paired with aging or priority classes. | Part 7 |
| S-LoRA | Serving thousands of LoRA adapters over one base with paged adapter memory and a unified batching kernel. | Part 17 |
| Sliding window attention | Attending only to the last W tokens, which bounds KV memory instead of letting it grow with the sequence. | Part 15 |
| SnapKV | Compressing the prefix by keeping the positions an observation window finds important; 3.6× decode at 16K with 8.2× memory efficiency. | Part 15 |
| Sparse attention | Attending to a selected subset of positions rather than all of them; the family that Quest and DuoAttention belong to. | Part 15 |
| Speculative decoding | Drafting several tokens cheaply and verifying them in one target-model pass, so more than one token emerges per step. | Part 11 |
| SRPT | Shortest-remaining-processing-time: SJF with preemption, reordering as remaining work changes. | Part 7 |
| Starvation | A request that never gets served because it is always outranked; the failure mode every scheduling policy must answer for. | Part 7 |
| Step time | Wall time for one engine iteration; its reciprocal is the per-sequence token rate, and it is set by the slower of bandwidth and compute. | Part 1 |
| StreamingLLM | Keeping a few attention-sink tokens plus a recent window; stable to millions of tokens without fine-tuning. | Part 15 |
| Structured output | Constraining generation to a grammar, regex or schema so the output is valid by construction. | Part 16 |
| Swap | Moving preempted KV to host DRAM instead of recomputing it; faster to restore, but it consumes host bandwidth and memory. | Part 5 |
| T | ||
| Tail latency | The slow end of the distribution, P90 and P99; the part of latency that SLOs and users actually experience. | Part 3 |
| Tensor parallelism (TP) | Splitting weights and compute across devices within a fast interconnect, paying an all-reduce every layer. | Part 12 |
| Thinking tokens | Long reasoning traces that inflate output tokens and therefore cost; the workload shape that breaks per-token pricing intuitions. | Part 18 |
| Throughput | Tokens or requests per second across all users; not the same as goodput, and not the same as a good user experience. | Part 3 |
| Tiered KV cache | Storing KV across HBM, DRAM, NVMe and remote memory with promotion and demotion between tiers. | Part 8 |
| Time-to-first-token (TTFT) | Arrival to first token, dominated by queueing and prefill; the number a streaming user perceives as responsiveness. | Part 3 |
| Token bucket | A rate limiter that refills at a fixed rate and rejects overflow; the per-tenant admission valve. | Part 17 |
| Token mask | The per-state vector of allowed next tokens; computing it cheaply is the whole game in structured decoding. | Part 16 |
| Tool call | A structured model output that invokes an external function, whose result returns to the context and extends the turn count. | Part 18 |
| TPOT | Time per output token after the first; the decode-speed metric that, with TTFT, defines a serving SLO. | Part 3 |
| Tree attention | Attention over a candidate tree that shares prefixes, letting all speculative branches be verified in one pass. | Part 11 |
| U | ||
| Unconstrained decoding | Free generation with no grammar; the baseline that the constraint tax is measured against. | Part 16 |
| V | ||
| Verifier | The target-model pass that accepts or rejects draft tokens and produces the accepted prefix. | Part 11 |
| Virtual token counter | A fair-share accounting scheme that charges tenants by tokens actually consumed, so one cannot monopolise the fleet. | Part 17 |
| Vocab mask | The full-vocabulary boolean mask applied each constrained step; its width makes tiny per-token overheads measurable. | Part 16 |
| W | ||
| Wait queue | Where admitted-but-not-running and preempted requests wait; its depth is the leading indicator of TTFT. | Part 7 |
| Warmup | Running throwaway inference passes to allocate memory and trigger kernel compilation before serving real traffic. | Part 19 |
| Weight load | Reading model weights from disk or network; the largest single term in the 40–90 s cold-start breakdown. | Part 19 |
| Weight-only quantization | Quantizing weights while leaving activations in FP16; the safest large win, and the one that fades above batch ~128. | Part 10 |
| Wide EP | Scaling expert parallelism across nodes and racks, as in DeepSeek's EP144 decode units, to spread hot experts. | Part 14 |
| Workload shape | The prompt, output and arrival mix that determines which bottleneck bites; chat, reasoning, agents and voice have different shapes. | Part 18 |
| X | ||
| XGrammar | Fast grammar-constrained decoding, reporting up to 100× faster mask generation than prior solutions with near-zero end-to-end overhead when overlapped. | Part 16 |
Quantization formats
Bytes per parameter and the accuracy bill
The table below is the series' shared reference for what each format buys and costs. The important column is not bits but bytes per parameter, because that is what a decode step reads; and the second important column is where the damage lands. Weight-only INT4 is usually benign; INT4 activations and INT4 KV need their own evaluation. The bytesPerParam column counts weights only and ignores scales, zero-points and KV, so treat it as a lower bound on the real footprint.
| Format | Bits | Bytes/param | Typical accuracy delta | Evaluation |
|---|
Engine feature matrix
What each server actually supports first-class
These four engines dominate open serving, and the differences that matter in production are the ones that are not on the marketing page: whether KV is paged, whether prefill and decode can be split, and how speculative decoding is plumbed. A tick means documented first-class support, not that every combination is production-hardened. The notes column carries the caveat that matters for each.
| Engine | Paged attn | Cont. batching | Prefix cache | Speculative | LoRA | Disaggregation | Notes |
|---|
Pricing reference
Historical list prices, labelled as such
The shape of the table is the lesson: cached input is roughly a tenth of fresh input, output is several times input, and the spread between a small and a frontier model is more than an order of magnitude. Any capacity plan that ignores the cache-hit column will overstate cost by roughly the hit rate.
| Model | Input / 1M | Output / 1M | Cached input / 1M | Note |
|---|
Further reading
Primary sources behind the numbers
- Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023 — the vLLM paper behind Part 5.
- Zheng et al., SGLang: Efficient Execution of Structured Language Model Programs, 2023 — RadixAttention and prefix caching in Part 8.
- Zhong et al., DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving, OSDI 2024 — the disaggregation result in Part 13.
- Qin et al., Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving, 2024 — the KVCache store referenced in Part 8.
- DeepSeek-AI, DeepSeek-V3 Technical Report, 2024 — MLA, FP8 and the MoE configuration in Parts 4, 10 and 14.
- Dong et al., XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models, MLSys 2025 — constrained decoding in Part 16.
- NVIDIA, Reducing Cold Start Latency with Model Streamer, 2025 — the 43.7 s → 7.5 s weight-load benchmark in Part 19.