Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Open loop, closed loop, and the mistake everyone makes

The load generator is part of the result

An open-loop harness submits requests on a fixed schedule — Poisson or constant rate — whether or not the server is keeping up. Arrival rate is an input; latency and queue depth are outputs. A closed-loop harness keeps a fixed number of clients, each of which sends its next request only after the previous one returns. There, concurrency is the input and the request rate is an output, governed by Little's law: throughput equals clients divided by (think time + latency).

The mistake is reading a closed-loop result as a property of the system. When the server slows down, a closed-loop client slows down with it, so the queue that produces a real p99 never forms: the harness reports a tail several times better than the machine can actually produce. Worse, the throughput plateau it reports is its own ceiling — clients divided by cycle time — not the server's capacity. That is why wrk-style constant-concurrency tests are useful for finding a bottleneck but useless for arguing about overload, and why incident reviews should always ask which harness produced the number. The training guide's inference chapter built the mechanisms — the KV cache, prefill versus decode, continuous batching — from the model's point of view; this part is the instrument that measures the system those mechanisms produce. The same tension appears in evaluation, where a benchmark score is a protocol as much as a model property; a serving benchmark is that idea with a queue attached.

Latency distribution for the same service under both harnesses. The selected harness is solid; the other is faded. The dashed line is the server's true capacity.

⚠️ The mistake: a closed-loop harness cannot build the queue that creates the tail, so its p99 is optimistic and its throughput ceiling is a harness constant. Quote the harness, the think time and the concurrency alongside every number, or the number is not reproducible.
2

The three harnesses you will actually run

vLLM bench, GenAI-Perf, MLPerf Inference

vllm bench is the engine's own tool: vllm bench serve drives a running OpenAI-compatible endpoint, while vllm bench throughput and vllm bench latency run the engine in-process. It is open-loop by default and reports TTFT, TPOT, ITL and E2E percentiles; it is the fastest way to answer "did my change help on this model".

GenAI-Perf (NVIDIA, a wrapper over the Triton perf analyzer) is a closed/open-loop capable client aimed at production endpoints, with per-request metrics, concurrency sweeps, and tokeniser-aware input/output token accounting. It is the natural fit when the deployment is TensorRT-LLM or Triton.

MLPerf Inference is the opposite end: a fixed, audited benchmark with prescribed models, datasets, scenario modes (Offline, Server, SingleStream, MultiStream) and accuracy constraints. Its value is not that it maps to your workload — it does not — but that it is comparable: identical load generation and rules across vendors. Use it for procurement and public claims; use the other two for your own regressions.

Real traces beat synthetic ones. Synthetic prompts are too short, too uniform and too independent. Production traces carry the properties that decide capacity: a heavy prompt-length tail, repeated prefixes that make a cache work, arrival bursts, and multi-turn sessions whose turns share context. A benchmark on 512-token independent prompts will understate prefill cost, overstate achievable batch, and miss the cache entirely. Replay a sampled, redacted trace whenever you can, and treat synthetic runs as unit tests rather than capacity evidence.

💡 Pick the harness for the question. Engine regression → vllm bench. Endpoint behaviour under concurrency → GenAI-Perf. Comparable public number → MLPerf Inference. Your own capacity plan → a replayed trace at your arrival process.
3

Warmup, and what must ride alongside a throughput number

A throughput number alone is not a measurement

The first requests a worker serves pay for CUDA context creation, CUDA-graph capture, allocator growth and kernel autotuning. A run that starts measuring immediately reports the transient: low early throughput, a fat early tail, and a mean that depends on how long you ran. Always discard a warmup period — enough requests to reach a stable step time, typically tens of seconds or a few hundred requests — and then measure a steady-state window. Report the warmup length too, because a number taken over a 10-second window is mostly warmup.

A defensible throughput or latency claim needs five things next to it: the input and output token distributions (not just their means), the arrival process (open-loop rate, or closed-loop concurrency and think time), the SLO the number is judged against, the software stack (engine, version, quantization, parallelism, scheduler flags), and the measurement window after warmup. Drop any one and the number is not comparable to anything — including your own next run.

💡 Write the card before the graph. "1,240 output tok/s/GPU, p99 TTFT 1.8 s, Qwen-class 7B FP8, 4,000/500-token trace, 2.0 rps open loop, vllm bench serve, 120 s window after 30 s warmup" is a measurement. "Fast" is not.
4

Why temperature 0 is not deterministic

Batch non-invariance, in real float32

Greedy decoding picks the argmax logit, so it looks reproducible. It is not, and the reason is arithmetic rather than sampling: float32 addition is not associative. A dot product or a reduction is computed by splitting the work across lanes and combining partial sums, and the split depends on the batch shape, the sequence length and the kernel's tiling. Change any of those and every partial rounds differently. The final logits differ by a few ULPs — enough to flip an argmax when two candidates are near-tied. The same request batched with different neighbours can emit a different token, which is why "temperature 0, seed fixed" still fails reproducibility tests. The numerics part covers where precision is lost; this demo shows that even exact FP32 loses it.

The demo below does the arithmetic for real, with Math.fround applied to every partial (a float32 rounding, not double). The left column is a sequential accumulation; the right is a pairwise tree. Toggle batch-invariant kernels to force one canonical order regardless of batch shape.

One float32 sum, two reduction splits. Bits from the top (31) to the bottom (0); the first differing bit is outlined.

Two near-tied candidate logits from the same arithmetic. The selected token changes with the reduction split unless kernels are batch-invariant.

💡 The fix is architectural: batch-invariant kernels fix the reduction order — and often the GPU-graph/atomic choices — so that a token's logits do not depend on its batch neighbours. It costs some throughput and is a requirement for reproducible evaluation, not the default.
5

The eight signals worth a dashboard

One row per failure mode

A GPU-utilisation gauge is not a dashboard. Utilisation can read 100% in a healthy saturated server and 100% in a server spinning on preemption. What distinguishes states is a small set of quantities that move independently. These eight cover the failures in this series; each maps to a different remedy.

1 · Queue depth and wait time. Requests admitted vs waiting, plus scheduler wait before first prefill. The earliest and cleanest overload signal.
2 · TTFT p50 and p99. p50 moves with prompt length and prefix-cache hit rate; a p99 that diverges from p50 is the queue, not the kernels.
3 · TPOT/ITL p50 and p99. Steady-state decode health. A p99 that spikes periodically is chunked-prefill interleaving or preemption churn.
4 · Goodput fraction. Share of admitted requests meeting both SLOs. Throughput can be flat while goodput collapses.
5 · KV-cache occupancy and preemptions. Used blocks, evictions, swaps and recomputes per second. Rising preemptions mean the cache, not the FLOPS, is the wall.
6 · Prefix-cache hit rate. Hit tokens over requested tokens. A sudden drop means the traffic's shared structure changed or routing scattered it.
7 · Batch occupancy and step time. Running sequences vs capacity, and milliseconds per iteration. It separates "slow model" from "empty batch".
8 · Adapter/tenant mix and error-retry rate. LoRA adapter count, per-tenant token rate, and 429/5xx/retry volume. Retries multiply load exactly when the system is saturated.
6

Incident signatures

Diagnose from the shape, not the spike

The dashboard below runs the shared discrete-event simulator in ServingSim.run. Press a lever to inject a failure and watch which of the eight signals move; the readout names the signature. The point is that each incident has a characteristic pattern across independent signals, and that pattern is what points at the fix. The long-prompt case below is the serving-side twin of long-context serving, where the KV cache rather than the queue is the limiting resource.

Per-request Gantt for the current run: grey queued, light prefill, dark decode, gold preempted. Inject a failure and watch the bands change.

SignatureWhat movesFirst check
Overload / queue saturationQueue depth ↑, TTFT p99 ↑↑, goodput ↓, throughput flatAdmission control and autoscaling headroom
Capacity lossPreemptions ↑, TPOT p99 ↑, KV occupancy ↑, batch occupancy ↓Replica health, then swap-vs-recompute policy
Long-prompt head-of-lineTTFT p50 ↑↑, ITL p99 spikes, short requests starveChunked prefill budget, then queueing policy
Prefix-cache thrashHit rate → 0, prefill volume ↑, KV pressure ↑, TTFT ↑Routing affinity and cache uniqueness
Retry stormError rate ↑, arrivals ↑, everything above at onceClient timeouts and backoff, not the model
7

Routing, cascades, semantic caching and the batch tier

Spend where the difficulty is

A router sends easy requests to a cheap model and hard ones to an expensive one; a cascade sends everything cheap first and escalates when a confidence check fails. Both trace an accuracy-versus-cost frontier: as the routing threshold moves, you buy accuracy with money along a curve whose shape is set by the difficulty distribution. The classic failure is a cascade that escalates on a noisy signal — you pay both models on the same request, and cost rises while accuracy stalls.

Semantic caching answers a request from a near-duplicate already served, at near-zero model cost. It shifts the whole frontier toward lower cost, but it can also cap accuracy: a cached answer is only as good as the match. Cache the stable, high-frequency, low-risk requests, and make the similarity threshold a tunable so you can trade cost against a measured accuracy drop.

The batch tier is the last routing decision: non-interactive work — offline scoring, evaluation, embeddings, bulk summarisation — belongs on a queue with no latency SLO, where the scheduler can use every token of batch capacity. Moving even a modest share of traffic off the interactive pool is often the cheapest capacity you can buy, and it is invisible in a token price. The scheduler part owns the queueing mechanics; the frontier below owns the trade.

Accuracy vs cost. The faded curve is no cache; the solid curve is the current semantic-cache hit rate. The dot is the current routing threshold.

8

RPS → GPUs → dollars, and cost per task

The closing planner

Capacity planning is a chain: a workload in requests per second and token lengths becomes GPU-seconds per second (prefill and decode separately, because one is compute-bound and one is memory-bound), divided by a utilisation target; GPU count times an hourly price becomes cost per million tokens; and that becomes cost per task only after you know how many tokens a task consumes and how often it succeeds. Per-token prices hide the last step, which is the one users feel. A task's token count is itself variable: reasoning models emit long thinking traces before any answer, and agents spend several turns over a shared context. Pricing by the token averages those away.

Every lever the series introduced is a toggle here. GQA and MLA change the KV footprint (the mechanism is in the training guide's architecture chapter); KV and weight quantization shrink the bytes read per step; prefix caching removes prefill work; chunked prefill and disaggregation change how the two phases share a pool; speculation buys decode tokens; MoE serves only the active parameters. Each lever is annotated with the GPUs it is currently saving by itself, and the ranked list below the planner sorts them for the current inputs, so the biggest lever changes when you change the workload. The model is a first-order planner, not a substitute for a measured run — but the ranking is what you would spend the next benchmark on.

GPUs saved by each lever in isolation, at the current inputs, sorted. Longer bar = bigger lever.

Cheat sheet

ConcernRule
Open vs closed loopOpen loop: rate in, latency out. Closed loop: concurrency in, rate out, and the queue never builds. Quote the harness, concurrency and think time.
Harness choicevllm bench for engine regressions · GenAI-Perf for endpoint concurrency · MLPerf Inference for comparable public claims.
Workload realismSynthetic prompts understate prefill cost and overstate cache hit rate. Replay a sampled trace for capacity.
WarmupDiscard tens of seconds or a few hundred requests; measure a steady-state window; report both.
Throughput claimReport token distributions, arrival process, SLO, stack/version and window. Anything less is not comparable.
DeterminismFP32 addition is not associative; batch shape changes reduction order and can flip an argmax at temperature 0. Batch-invariant kernels fix it.
Eight signalsQueue wait · TTFT p50/p99 · TPOT p50/p99 · goodput · KV occupancy/preemptions · prefix hit rate · batch occupancy/step time · tenant mix/retries.
Incident triageQueue growth = overload · preemptions+TPOT = capacity loss · TTFT p50 + ITL spike = long prompt · hit rate → 0 = cache thrash · errors+arrivals = retry storm.
CostCost per million tokens falls with throughput and utilisation; cost per successful task also moves with turns and retries. Optimise the task, not the token.

Further reading

Cheat sheet

InstrumentWhat it capturesRule
Open-loop harnessFixed arrival rate; latency is the outputUse for overload and tail claims
Closed-loop harnessFixed clients; rate = clients ÷ (think + latency)Optimistic p99; the plateau is the harness ceiling
WarmupCUDA context, graph capture, autotuningDiscard it, and report the window you measured
Measurement cardToken distributions, arrival process, SLO, stack, windowWithout all five, a number isn’t comparable
Batch invarianceFloat addition isn’t associativeGreedy decoding isn’t reproducible across batch shapes
Capacity planRPS → GPU-seconds per request → GPUs → USD per taskPlan from a replayed trace, not averages
?

Check your understanding

0/5 answered