Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Accelerators: the bandwidth table

Capacity, bandwidth, FLOPs, fabric, power

A serving system is bounded below by the memory bandwidth of the chip it runs on. HBM capacity decides how many KV blocks fit and therefore how many sequences can be resident; bandwidth decides how fast a decode step can sweep the weights and the cache; peak FLOPs set the prefill ceiling; the interconnect decides whether you may shard at all. Read the table by asking which column is smallest for your workload. A single-stream decode job is almost always bandwidth-bound, so bwTBs is the number that predicts tokens per second; a long-prompt prefill job is compute-bound, so bf16TFLOPs and fp8TFLOPs matter more; a multi-node tensor-parallel deployment lives or dies on the fabric column.

The NVIDIA headline FLOPs in the upstream source include 2:4 sparsity, so the dense values rendered here are half the press-release number. AMD and Google publish dense figures directly. A dash means the format is not natively supported by that part, which is why FP4 is a Blackwell story and FP8 is a Hopper-and-later story.

AcceleratorHBMBWBF16FP8FP4InterconnectTDP

2

Models and their KV configurations

The two numbers that drive a capacity plan

Model parameters set the weight bytes a decode step must read; KV bytes per token set how many sequences fit beside those weights. Both are computed from public configs, not estimated. For an MHA or GQA model the per-token cache is 2 × layers × n_kv_heads × d_head × 2 bytes — the leading 2 for K and V, the trailing 2 for FP16. MLA breaks the formula because it caches a single low-rank latent plus a decoupled RoPE key, and SWA-hybrid models bound the cache at the sliding-window length instead of letting it grow linearly, so read their rows as ceilings rather than rates. The KV bytes column is the one the concurrency calculators in the series use; multiply by context length and concurrency to see whether a model fits.

ModelLayersQ headsKV headsd_headSchemeKV per token (FP16)Params total / activeContext

3

Back-of-envelope formulas

Six expressions that explain most serving behaviour

Every interactive demo in the series is a version of one of these. Keep them in a notebook; they are accurate to a factor you can reason about, which is the point. Bytes per element is written β and equals 2 for BF16/FP16, 1 for FP8/INT8 and 0.5 for INT4.

KV bytes

$$\text{KV bytes} \;=\; 2\,L\,H_{kv}\,d_{head}\,T\,b\,\beta$$

L layers, Hkv KV heads, head width dhead, T cached tokens, batch b. For MLA replace 2 Hkv dhead with the cached latent width (576 for DeepSeek-V3).

Step time

$$T_{\text{step}} \;\approx\; \max\!\left(\frac{P\,b_p + \text{KV}_{\text{read}}}{\text{BW}},\; \frac{\text{FLOPs}}{\text{Peak}}\right) + T_{\text{overhead}}$$

The max of the memory-bound and compute-bound terms, plus launch and synchronization overhead. Decode at small batch is memory-bound; prefill is compute-bound.

Expected speculative tokens

$$\mathbb{E}[\text{tokens}] \;=\; \frac{1-\alpha^{\gamma+1}}{1-\alpha}$$

Per draft round with acceptance rate α and γ draft tokens: 1 + α + α2 + … + αγ. At α=0.8 and γ=4 the round yields about 3.36 tokens.

Pipeline bubble fraction

$$\text{bubble} \;=\; \frac{p-1}{m+p-1}$$

p pipeline stages and m microbatches. The fraction of stage-time spent idle while the pipe fills and drains; it vanishes only as m grows large.

All-reduce bytes

$$\text{bytes/GPU} \;=\; 2\cdot\frac{N-1}{N}\cdot S$$

Ring all-reduce moving message size S across N ranks. The factor 2 is the reduce-scatter plus all-gather; it is the per-step tax every TP group pays.

Goodput

$$G \;=\; \frac{1}{T}\sum_{r} \mathbb{1}\!\left[\text{TTFT}_r \le \text{SLO}_{\text{ttft}} \;\wedge\; \text{TPOT}_r \le \text{SLO}_{\text{tpot}}\right]$$

Requests per second that satisfy both latency limits. Raw throughput counts work you cannot bill for; goodput counts only the work that met its contract.

4

Glossary

Every term, linked to the part that introduces it

💡 Filter the list. Type any fragment — a term, a number, a technology — and the table hides non-matching rows and reports how many remain. Matching is case-insensitive and looks at both the term and its definition.
TermWhat it meansIntroduced in
A
Acceptance rateExpected fraction of speculative draft tokens the target model accepts in a verification pass; the single number that decides whether speculation pays.Part 11
Accelerator zooThe GPU/TPU/Trainium landscape compared by HBM, bandwidth, FLOPs and interconnect rather than by brand.Part 2
Activation outlierA few input channels whose values dwarf the rest, setting a coarse step for every ordinary channel under per-tensor scaling.Part 10
Active parametersParameters actually touched per token in a mixture-of-experts model (37B of 671B for DeepSeek-V3), as distinct from the total resident.Part 14
Admission controlDeciding whether to accept a request into the running set or shed it, based on KV headroom, capacity and the SLO.Part 7
AgingIncreasing a waiting request's effective priority over time so that long jobs cannot starve behind a stream of short ones.Part 7
All-reduceSumming a tensor across ranks and broadcasting the result; the synchronization every tensor-parallel step pays.Part 12
All-to-allEvery rank exchanges a different slice with every other rank; the dispatch/combine primitive of expert parallelism.Part 14
Arithmetic intensityFLOPs performed per byte moved; about 1–2 for a decode step and far higher for prefill, which is why the two live in different regimes of the roofline.Part 1
Attention sinkThe first few tokens that absorb a disproportionate share of attention and must be kept for stability when the cache is trimmed.Part 15
B
BackpressureReturning a 429 or slowing the client instead of letting an unbounded queue convert overload into an SLO breach for everyone.Part 7
Batch invarianceWhether a request's output is bit-identical regardless of the other requests it was batched with; it is not, because floating-point addition is not associative.Part 20
Batch size (max_num_seqs)How many sequences the engine runs per iteration; the knob that moves a step from memory-bound toward compute-bound.Part 6
Block managerThe allocator mapping a sequence's logical token positions to physical KV blocks, and the component preemption calls into.Part 5
Block sizeTokens per KV page, conventionally 16; the granularity at which internal fragmentation is paid.Part 5
Block tableThe per-sequence array that translates logical block indices to physical block indices, so paged KV needs no contiguity.Part 5
Bottleneck (memory vs compute)Which resource limits a step; decode is memory-bandwidth-bound, prefill is compute-bound, and batch size moves the crossing point.Part 1
C
Cache hit rateFraction of prompt tokens whose KV was already resident; the statistic that turns prefix caching from an optimisation into a throughput multiplier.Part 8
Cache-aware routingSending a request to the replica that already holds its prefix instead of the least-loaded one.Part 8
CalibrationRunning representative data through the model to measure activation ranges before quantizing; the ranges are data-dependent, so this is not really post-training.Part 10
Capacity planningConverting an RPS and latency target into GPUs and dollars, including the queueing headroom the mean hides.Part 20
Chunked prefillSplitting a long prompt's prefill across iterations so decode requests are not stalled for the whole prompt's compute time.Part 6
Cold startThe 40–90 s from pod schedule to first token: image pull, process init, weight load, graph capture and warmup.Part 19
Compute-boundStep time set by FLOPs rather than bytes; the high-batch and long-prefill regime.Part 1
Connector (KV)The transport that ships a finished prefill's KV blocks to a decode worker, often over RDMA or NIXL.Part 13
Context lengthTokens the model can attend to; KV memory and decode cost grow linearly with it while prefill cost grows quadratically.Part 4
Context parallelism (CP)Splitting the sequence dimension across devices so a long context's attention and KV fit, with activations passed around the ring.Part 12
Continuous batchingAdmitting and retiring sequences every iteration rather than per fixed batch, so a slow request never idles a whole batch.Part 6
Copy-on-writeTwo sequences share a physical KV block until one writes, at which point the block is copied; how shared prefixes are stored once.Part 5
Cost per taskDollars per completed task rather than per token; the number a budget actually sees once retries and thinking tokens are counted.Part 20
CUDA graphA captured sequence of kernel launches replayed as a single operation to remove per-launch CPU overhead from a decode step.Part 9
D
Data parallelism (DP)Replicating the model and splitting requests across copies; the default until weights no longer fit one device.Part 12
DecodeGenerating one token at a time, memory-bandwidth-bound because every step sweeps the weights through HBM.Part 1
Decode workerThe pool in a disaggregated deployment that holds KV and generates tokens after prefill has produced it.Part 13
DeepEPDeepSeek's open-source MoE communication library implementing grouped all-to-all dispatch and combine.Part 14
DefragmentationReclaiming scattered free KV blocks, or moving pages so there is no need to; paging turns this from a stop-the-world op into an allocation detail.Part 5
Disaggregation (PD)Running prefill and decode in separate pools, each sized for its own bottleneck, with KV transferred between them.Part 13
Distributed KV cachePartitioning a long context's KV across devices, for example with ring attention, so no single GPU holds all of it.Part 15
DP-attentionCombining data parallelism with attention sharding so KV is not replicated across replicas, trading communication for capacity.Part 12
Draft modelThe small model that proposes speculative tokens for the target model to verify; its agreement rate sets the speedup.Part 11
Draft treeSeveral candidate continuations explored in parallel and verified together, as in EAGLE and Medusa.Part 11
DuoAttentionKeeping full KV for a small set of retrieval heads and constant-length KV for streaming heads; up to 2.18× decode speedup for MHA.Part 15
E
EAGLESpeculative decoding that drafts in the target model's feature space; EAGLE-2 reports 3.05–4.26× and EAGLE-3 up to 6.5×.Part 11
End-to-end latencyArrival to last token: queueing plus prefill plus decode, and the number a streaming UI hides behind TTFT.Part 3
Energy per tokenJoules per generated token, roughly power × step time divided by tokens per step; the metric that makes rack power a serving variable.Part 2
EngineThe serving software — vLLM, SGLang, TensorRT-LLM, llama.cpp — that schedules requests and runs the kernels.Part 9
EPLBExpert-parallel load balancer: replicates hot experts onto less-loaded ranks to even out MoE traffic; roughly a 20% recovery is cited for DeepSeek-V3-class deployments.Part 14
Eviction (KV)Dropping low-value KV entries to cap cache size, whether by attention score, recency or query relevance.Part 15
ExpertOne of many parallel MLPs that a router selects per token; the unit sharded by expert parallelism.Part 14
Expert parallelism (EP)Sharding experts across devices, one or more per rank, and paying an all-to-all to dispatch tokens and combine results.Part 14
Expert routerThe gate that scores a token and picks the top-k experts to handle it; its load distribution is what makes hot experts a problem.Part 14
F
Fair queueingScheduling that gives each tenant a bounded share under contention, so one workload cannot consume the fleet.Part 17
FCFSFirst-come-first-served ordering; simple and starvation-free, but one long request blocks everything behind it.Part 7
First-token latencyAnother name for TTFT, usefully split into queueing time and prefill time when diagnosing a regression.Part 3
FlashAttentionIO-aware attention that tiles Q, K and V through SRAM instead of materializing the full attention matrix in HBM.Part 9
FP44-bit floating point, native on Blackwell at 0.5 bytes per parameter; the format that makes B200's FP4 peak a headline.Part 10
FP8 (E4M3 / E5M2)8-bit floating point, native on Hopper and Blackwell; E4M3 with per-block scaling is DeepSeek-V3's serving format.Part 10
FragmentationWasted KV memory from rounding a sequence up to a block boundary or from gaps between reserved regions.Part 5
FSM (finite state machine)The automaton a regular expression compiles to; its current state indexes the set of legal next tokens.Part 16
G
GatewayThe front door that authenticates, rate-limits, observes and routes requests before any GPU sees them.Part 19
GoodputRequests per second that meet their latency SLOs; raw throughput counts work you cannot bill for.Part 3
GPU utilizationFraction of time SMs are busy; routinely near 100% while goodput is poor, which is why it is a lie as a serving signal.Part 19
Graph captureRecording a decode iteration's kernel launches into a CUDA graph so replay costs one launch instead of dozens.Part 9
Graph reuseCaching captured CUDA graphs across process starts to remove graph-capture time from a cold start.Part 19
GQAGrouped-query attention: several query heads share one KV head, cutting KV by 8× at Llama-3-70B geometry with little quality loss.Part 4
H
H2OKV eviction that keeps a small set of heavy-hitter tokens plus recent ones; up to 29× throughput over naive baselines.Part 15
HBMHigh-bandwidth memory: its capacity caps concurrency, its bandwidth caps decode speed, and both are read from the accelerator table.Part 2
Head dimensionWidth of one attention head's vectors; it multiplies directly into KV bytes per token.Part 4
Head-of-line blockingOne long request delaying all the shorter ones queued behind it, the failure FCFS invites.Part 7
Heavy hitterA token that receives disproportionate attention and is therefore worth keeping when the cache must shrink.Part 15
Hidden stateThe residual-stream vector each layer reads and writes; the tensor whose size the weights matrix multiplies per token.Part 1
Hot expertAn expert that receives far more tokens than average; the top few percent of experts absorb a large share of traffic, which EPLB targets.Part 14
I
ICIThe TPU interconnect; Google's analogue of NVLink and the fabric that decides whether a TPU shard communicates cheaply.Part 2
Image tokensVision input expanded into context tokens; large, prefill-heavy and a reason multimodal workloads invert the prefill/decode ratio.Part 18
In-flight batchingTensorRT-LLM's name for continuous batching: sequences enter and leave the running set mid-flight.Part 9
InterconnectThe link (NVLink, NVSwitch, ICI, NeuronLink) that determines whether tensor and expert parallelism pay for themselves.Part 2
Inter-token latency (ITL)The gap between consecutive output tokens; the inverse of TPOT and the number a chat user actually perceives as speed.Part 3
J
JSON schemaThe grammar that constrains generation to a valid structure; the commonest reason structured decoding is deployed.Part 16
Jump-forward decodingEmitting a run of forced tokens at once when the grammar admits only one continuation, rather than one step at a time.Part 16
K
KernelA compiled GPU routine; the unit whose launch overhead CUDA graphs remove and whose shape determines achieved efficiency.Part 9
Kubernetes (K8s)The orchestrator that schedules serving pods, runs probes and drives autoscaling, and therefore sets the cold-start floor.Part 19
KV blockA fixed-size chunk of KV cache, the unit PagedAttention allocates and the block table maps.Part 5
KV cacheCached keys and values so decode reuses attention state instead of recomputing it; the tensor that caps concurrency.Part 4
KV cache hitA prefix whose KV was already computed and can be reused; a hit converts prefill work into a memory read.Part 8
KV compressionStoring or summarizing fewer, cheaper KV entries, whether by quantization, eviction or low-rank projection.Part 15
KV offloadMoving KV to a lower tier (DRAM or NVMe) and fetching it back when needed, trading latency for resident capacity.Part 8
KV quantizationStoring KV in FP8, INT8 or INT4 to fit more sequences; every error is re-read on every decode step, so evaluate, do not assume.Part 10
KV transferShipping KV blocks over RDMA or NIXL from a prefill worker to a decode worker; the cost that bounds disaggregation.Part 13
KV-aware routingLoad balancing by which replica holds the request's prefix, not merely by which has spare capacity.Part 8
L
Latency SLOThe maximum TTFT or TPOT a request may have and still count as served; goodput is measured against it.Part 3
Length predictionGuessing an output's length to schedule and admit better; wrong guesses either strand capacity or miss deadlines.Part 7
LLGuidanceA constrained-decoding library that compiles grammars into token masks for fast structured generation.Part 16
Load balancerDistributes requests across replicas; a prefix-aware balancer turns a cache hit into a cheaper request.Part 19
Load sheddingRejecting work deliberately to protect the requests already admitted; the honest alternative to an unbounded queue.Part 7
LogitThe raw pre-softmax score for each vocabulary entry; sampling turns the logit vector into one token.Part 1
Long-contextServing very long prompts: quadratic prefill work, linear KV, and the eviction and routing machinery that makes it affordable.Part 15
LoRAA low-rank adapter over frozen base weights; serving many of them over one base is the multi-tenant pattern.Part 17
M
Mask (logit)Setting disallowed tokens to −∞ before sampling so the grammar is never violated.Part 16
MBUMemory-bandwidth utilization: achieved bytes per second over peak; the utilisation that actually predicts decode speed.Part 2
MedusaAdding decoding heads to a backbone so several future tokens are predicted and verified in parallel; 2.3–3.6× reported.Part 11
Memory wallThe widening gap between compute growth and bandwidth growth that makes decode memory-bound rather than compute-bound.Part 1
MFUModel-FLOPs utilization: achieved FLOPs per second over the hardware peak; the prefill-side analogue of MBU.Part 2
MLAMulti-head latent attention: DeepSeek's low-rank KV that caches one compressed latent plus a RoPE key, about 57× smaller than MHA at the same geometry.Part 4
Model StreamerNVIDIA Run:ai's parallel weight loader; its benchmark drops a 15 GB model's load from 43.7 s to 7.5 s as concurrency rises.Part 19
MoEMixture of experts: many parameters resident but only a few active per token, which decouples capacity from FLOPs.Part 14
MooncakeMoonshot's KVCache-centric disaggregated store; its vLLM integration reports 3.8× throughput and 46× lower P50 TTFT on agentic traces.Part 8
MQAMulti-query attention: a single KV head, up to 64× less cache than MHA at the cost of quality headroom.Part 4
Multi-LoRAServing many adapters over one base model in a single batch, with adapters paged in and out like KV blocks.Part 17
Multi-tenancySharing serving capacity across independent workloads or customers, with fairness and isolation rather than one global queue.Part 17
MXFP4Microscaling 4-bit with a shared block exponent; the weight format named by gpt-oss-120b and accelerated natively on Blackwell.Part 10
N
N-gram / prompt lookupSpeculating the continuation by copying matching n-grams from the prompt; no draft model, and the gain is entirely workload-dependent.Part 11
NIXLNVIDIA's inference transfer library for moving KV between disaggregated workers over RDMA and other fabrics.Part 13
Noisy neighbourA tenant whose load degrades the latency of others sharing the same replicas; the reason fair queueing is a serving feature.Part 17
Node (prefill / decode)The unit of disaggregated capacity; DeepSeek runs 32-GPU prefill units against 144-GPU decode units in its published split.Part 13
NVFP4Blackwell's 4-bit floating-point format adding a per-block FP8 scale on top of E2M1 values.Part 10
O
OOMOut-of-memory from KV growth; paging, preemption and admission control exist to prevent it from taking down a worker.Part 5
Open vs closed loopWhether the client backs off when latency rises (closed) or pushes regardless (open); benchmarks must state which they use.Part 20
OutlinesThe early FSM-indexing structured-generation library which introduced jump-forward decoding and near-zero masking overhead.Part 16
OverloadThe regime where arrivals exceed capacity; the regime admission control, load shedding and backpressure exist for.Part 7
P
Page (KV)A physical block of KV cache, the unit PagedAttention maps from a sequence's logical positions.Part 5
Paged KV cacheNon-contiguous, block-paged KV storage that removes the need to reserve a sequence's maximum length up front.Part 5
PagedAttentionThe vLLM mechanism that pages KV and translates logical to physical blocks with a block table.Part 5
PaddingWasted batch slots from forcing variable-length sequences into a rectangle; packing and paging avoid it.Part 6
ParallelismSplitting a model's compute across devices: tensor, pipeline, expert, data (attention) and context parallelism.Part 12
Percentile (P50 / P90 / P99)The latency below which that fraction of requests fall; the mean hides the tail that actually breaks SLOs.Part 3
Pipeline bubbleIdle time while pipeline stages fill and drain, proportional to (p−1)/(m+p−1) for p stages and m microbatches.Part 12
Pipeline parallelism (PP)Splitting model layers across stages and streaming microbatches through them; needs less bandwidth than TP but adds latency and bubbles.Part 12
PrefillProcessing the whole prompt in parallel; compute-bound, and the source of TTFT.Part 1
Prefill-decode ratioThe mix of prompt and output tokens in a workload; agentic and multimodal traffic invert the ratio that chat assumes.Part 18
Prefill workerThe pool in a disaggregated deployment that only computes prompt KV and hands it to decoders.Part 13
PreemptionEvicting a running sequence's KV to make room, then recomputing or swapping it back when capacity returns.Part 5
Prefix affinityThe routing preference for the replica that already holds a shared prefix, so a cache hit is not thrown away.Part 8
Prefix cachingReusing KV for shared prompt prefixes across requests; a cache hit turns prefill time into a memory read.Part 8
Priority schedulingOrdering the queue by request class rather than arrival, which needs aging or quotas to stay starvation-free.Part 7
Pushdown automatonA stack machine needed for nested grammars such as balanced JSON, since a finite-state machine cannot count depth.Part 16
PyramidKVLayer-wise pyramidal KV budgets, more retention in low layers and less in high ones; near-full quality at roughly 12% of KV.Part 15
Q
QuantizationStoring weights, activations or KV at lower precision to save bytes; in serving the win is bandwidth, not FLOPs.Part 10
Query-aware evictionChoosing KV entries to keep based on the current query rather than only recency or global attention mass.Part 15
QuestQuery-aware top-K KV page selection using per-page key ranges; up to 7.03× reported latency reduction.Part 15
R
RadixAttentionSGLang's radix-tree prefix cache; reported around 6× higher throughput on structured workloads.Part 8
Radix treeThe prefix tree that indexes cached KV by shared token sequences and makes longest-prefix matching cheap.Part 8
RecomputeRebuilding a preempted sequence's KV from its tokens instead of moving it; cheaper than a fetch below roughly 1,900 tokens.Part 5
Rejection samplingThe rule that accepts a draft token only when the target model's sample agrees, keeping the output distribution exact.Part 11
Request routingChoosing a replica for each request; the choice changes cache hit rate, queue depth and therefore latency.Part 19
Ring attentionDistributing a long sequence's attention across devices in a ring so each holds part of the KV and passes activations on.Part 15
RooflineThe model relating attainable FLOPs to arithmetic intensity, peak compute and peak bandwidth; it explains the memory wall.Part 1
RPSRequests per second; the input to capacity planning once each request's token profile is known.Part 20
S
SamplingTurning logits into a token with greedy, temperature or top-p rules; it is the last step of every decode iteration.Part 1
Scale-to-zeroRemoving the last replica and paying a cold start to return; honest only when the traffic is genuinely bursty.Part 19
SchedulerThe component deciding what to admit, preempt and run each iteration, under KV and capacity limits.Part 7
Semantic cachingReusing an answer when a new query is semantically equivalent to an old one; a workload-level cache above the KV cache.Part 18
Serving SLOThe latency and availability contract an API makes, expressed on TTFT, TPOT and error rate rather than on averages.Part 3
SGMVA batched LoRA kernel that applies many adapters in one pass by treating the batch as a segmented matrix-vector product.Part 17
ShardingSplitting weights, experts or KV across devices; every sharding choice trades memory for communication.Part 12
SJFShortest-job-first ordering; improves mean latency but starves long jobs unless paired with aging or priority classes.Part 7
S-LoRAServing thousands of LoRA adapters over one base with paged adapter memory and a unified batching kernel.Part 17
Sliding window attentionAttending only to the last W tokens, which bounds KV memory instead of letting it grow with the sequence.Part 15
SnapKVCompressing the prefix by keeping the positions an observation window finds important; 3.6× decode at 16K with 8.2× memory efficiency.Part 15
Sparse attentionAttending to a selected subset of positions rather than all of them; the family that Quest and DuoAttention belong to.Part 15
Speculative decodingDrafting several tokens cheaply and verifying them in one target-model pass, so more than one token emerges per step.Part 11
SRPTShortest-remaining-processing-time: SJF with preemption, reordering as remaining work changes.Part 7
StarvationA request that never gets served because it is always outranked; the failure mode every scheduling policy must answer for.Part 7
Step timeWall time for one engine iteration; its reciprocal is the per-sequence token rate, and it is set by the slower of bandwidth and compute.Part 1
StreamingLLMKeeping a few attention-sink tokens plus a recent window; stable to millions of tokens without fine-tuning.Part 15
Structured outputConstraining generation to a grammar, regex or schema so the output is valid by construction.Part 16
SwapMoving preempted KV to host DRAM instead of recomputing it; faster to restore, but it consumes host bandwidth and memory.Part 5
T
Tail latencyThe slow end of the distribution, P90 and P99; the part of latency that SLOs and users actually experience.Part 3
Tensor parallelism (TP)Splitting weights and compute across devices within a fast interconnect, paying an all-reduce every layer.Part 12
Thinking tokensLong reasoning traces that inflate output tokens and therefore cost; the workload shape that breaks per-token pricing intuitions.Part 18
ThroughputTokens or requests per second across all users; not the same as goodput, and not the same as a good user experience.Part 3
Tiered KV cacheStoring KV across HBM, DRAM, NVMe and remote memory with promotion and demotion between tiers.Part 8
Time-to-first-token (TTFT)Arrival to first token, dominated by queueing and prefill; the number a streaming user perceives as responsiveness.Part 3
Token bucketA rate limiter that refills at a fixed rate and rejects overflow; the per-tenant admission valve.Part 17
Token maskThe per-state vector of allowed next tokens; computing it cheaply is the whole game in structured decoding.Part 16
Tool callA structured model output that invokes an external function, whose result returns to the context and extends the turn count.Part 18
TPOTTime per output token after the first; the decode-speed metric that, with TTFT, defines a serving SLO.Part 3
Tree attentionAttention over a candidate tree that shares prefixes, letting all speculative branches be verified in one pass.Part 11
U
Unconstrained decodingFree generation with no grammar; the baseline that the constraint tax is measured against.Part 16
V
VerifierThe target-model pass that accepts or rejects draft tokens and produces the accepted prefix.Part 11
Virtual token counterA fair-share accounting scheme that charges tenants by tokens actually consumed, so one cannot monopolise the fleet.Part 17
Vocab maskThe full-vocabulary boolean mask applied each constrained step; its width makes tiny per-token overheads measurable.Part 16
W
Wait queueWhere admitted-but-not-running and preempted requests wait; its depth is the leading indicator of TTFT.Part 7
WarmupRunning throwaway inference passes to allocate memory and trigger kernel compilation before serving real traffic.Part 19
Weight loadReading model weights from disk or network; the largest single term in the 40–90 s cold-start breakdown.Part 19
Weight-only quantizationQuantizing weights while leaving activations in FP16; the safest large win, and the one that fades above batch ~128.Part 10
Wide EPScaling expert parallelism across nodes and racks, as in DeepSeek's EP144 decode units, to spread hot experts.Part 14
Workload shapeThe prompt, output and arrival mix that determines which bottleneck bites; chat, reasoning, agents and voice have different shapes.Part 18
X
XGrammarFast grammar-constrained decoding, reporting up to 100× faster mask generation than prior solutions with near-zero end-to-end overhead when overlapped.Part 16
5

Quantization formats

Bytes per parameter and the accuracy bill

The table below is the series' shared reference for what each format buys and costs. The important column is not bits but bytes per parameter, because that is what a decode step reads; and the second important column is where the damage lands. Weight-only INT4 is usually benign; INT4 activations and INT4 KV need their own evaluation. The bytesPerParam column counts weights only and ignores scales, zero-points and KV, so treat it as a lower bound on the real footprint.

FormatBitsBytes/paramTypical accuracy deltaEvaluation

6

Engine feature matrix

What each server actually supports first-class

These four engines dominate open serving, and the differences that matter in production are the ones that are not on the marketing page: whether KV is paged, whether prefill and decode can be split, and how speculative decoding is plumbed. A tick means documented first-class support, not that every combination is production-hardened. The notes column carries the caveat that matters for each.

EnginePaged attnCont. batchingPrefix cacheSpeculativeLoRADisaggregationNotes

7

Pricing reference

Historical list prices, labelled as such

⚠️ These prices are historical. They are early-to-mid 2025 vendor list prices and several of these models are now superseded or repriced. They are here as clearly labelled reference points for the cost arithmetic in Part 20, not as current quotes. Verify against the live vendor pages before relying on any number.

The shape of the table is the lesson: cached input is roughly a tenth of fresh input, output is several times input, and the spread between a small and a frontier model is more than an order of magnitude. Any capacity plan that ignores the cache-hit column will overstate cost by roughly the hit rate.

ModelInput / 1MOutput / 1MCached input / 1MNote

📚

Further reading

Primary sources behind the numbers