Benchmarking, cost and capacity planning
A serving number is not a property of the model. It is a property of the model, the hardware, the workload, the harness and the SLO, sampled over a window. This final part builds the five instruments that keep those apart: a harness that can tell overload from a client ceiling, a batch-invariance demo that shows why greedy decoding is not reproducible, a dashboard whose incidents have names, a router frontier, and a capacity planner that goes from requests per second to GPUs to dollars per task.
Open loop, closed loop, and the mistake everyone makes
The load generator is part of the result
An open-loop harness submits requests on a fixed schedule — Poisson or constant rate — whether or not the server is keeping up. Arrival rate is an input; latency and queue depth are outputs. A closed-loop harness keeps a fixed number of clients, each of which sends its next request only after the previous one returns. There, concurrency is the input and the request rate is an output, governed by Little's law: throughput equals clients divided by (think time + latency).
The mistake is reading a closed-loop result as a property of the system. When the server slows down, a closed-loop client slows down with it, so the queue that produces a real p99 never forms: the harness reports a tail several times better than the machine can actually produce. Worse, the throughput plateau it reports is its own ceiling — clients divided by cycle time — not the server's capacity. That is why wrk-style constant-concurrency tests are useful for finding a bottleneck but useless for arguing about overload, and why incident reviews should always ask which harness produced the number. The training guide's inference chapter built the mechanisms — the KV cache, prefill versus decode, continuous batching — from the model's point of view; this part is the instrument that measures the system those mechanisms produce. The same tension appears in evaluation, where a benchmark score is a protocol as much as a model property; a serving benchmark is that idea with a queue attached.
Latency distribution for the same service under both harnesses. The selected harness is solid; the other is faded. The dashed line is the server's true capacity.
The three harnesses you will actually run
vLLM bench, GenAI-Perf, MLPerf Inference
vllm bench is the engine's own tool: vllm bench serve drives a running OpenAI-compatible endpoint, while vllm bench throughput and vllm bench latency run the engine in-process. It is open-loop by default and reports TTFT, TPOT, ITL and E2E percentiles; it is the fastest way to answer "did my change help on this model".
GenAI-Perf (NVIDIA, a wrapper over the Triton perf analyzer) is a closed/open-loop capable client aimed at production endpoints, with per-request metrics, concurrency sweeps, and tokeniser-aware input/output token accounting. It is the natural fit when the deployment is TensorRT-LLM or Triton.
MLPerf Inference is the opposite end: a fixed, audited benchmark with prescribed models, datasets, scenario modes (Offline, Server, SingleStream, MultiStream) and accuracy constraints. Its value is not that it maps to your workload — it does not — but that it is comparable: identical load generation and rules across vendors. Use it for procurement and public claims; use the other two for your own regressions.
Real traces beat synthetic ones. Synthetic prompts are too short, too uniform and too independent. Production traces carry the properties that decide capacity: a heavy prompt-length tail, repeated prefixes that make a cache work, arrival bursts, and multi-turn sessions whose turns share context. A benchmark on 512-token independent prompts will understate prefill cost, overstate achievable batch, and miss the cache entirely. Replay a sampled, redacted trace whenever you can, and treat synthetic runs as unit tests rather than capacity evidence.
vllm bench. Endpoint behaviour under concurrency → GenAI-Perf. Comparable public number → MLPerf Inference. Your own capacity plan → a replayed trace at your arrival process.Warmup, and what must ride alongside a throughput number
A throughput number alone is not a measurement
The first requests a worker serves pay for CUDA context creation, CUDA-graph capture, allocator growth and kernel autotuning. A run that starts measuring immediately reports the transient: low early throughput, a fat early tail, and a mean that depends on how long you ran. Always discard a warmup period — enough requests to reach a stable step time, typically tens of seconds or a few hundred requests — and then measure a steady-state window. Report the warmup length too, because a number taken over a 10-second window is mostly warmup.
A defensible throughput or latency claim needs five things next to it: the input and output token distributions (not just their means), the arrival process (open-loop rate, or closed-loop concurrency and think time), the SLO the number is judged against, the software stack (engine, version, quantization, parallelism, scheduler flags), and the measurement window after warmup. Drop any one and the number is not comparable to anything — including your own next run.
vllm bench serve, 120 s window after 30 s warmup" is a measurement. "Fast" is not.Why temperature 0 is not deterministic
Batch non-invariance, in real float32
Greedy decoding picks the argmax logit, so it looks reproducible. It is not, and the reason is arithmetic rather than sampling: float32 addition is not associative. A dot product or a reduction is computed by splitting the work across lanes and combining partial sums, and the split depends on the batch shape, the sequence length and the kernel's tiling. Change any of those and every partial rounds differently. The final logits differ by a few ULPs — enough to flip an argmax when two candidates are near-tied. The same request batched with different neighbours can emit a different token, which is why "temperature 0, seed fixed" still fails reproducibility tests. The numerics part covers where precision is lost; this demo shows that even exact FP32 loses it.
The demo below does the arithmetic for real, with Math.fround applied to every partial (a float32 rounding, not double). The left column is a sequential accumulation; the right is a pairwise tree. Toggle batch-invariant kernels to force one canonical order regardless of batch shape.
One float32 sum, two reduction splits. Bits from the top (31) to the bottom (0); the first differing bit is outlined.
Two near-tied candidate logits from the same arithmetic. The selected token changes with the reduction split unless kernels are batch-invariant.
The eight signals worth a dashboard
One row per failure mode
A GPU-utilisation gauge is not a dashboard. Utilisation can read 100% in a healthy saturated server and 100% in a server spinning on preemption. What distinguishes states is a small set of quantities that move independently. These eight cover the failures in this series; each maps to a different remedy.
Incident signatures
Diagnose from the shape, not the spike
The dashboard below runs the shared discrete-event simulator in ServingSim.run. Press a lever to inject a failure and watch which of the eight signals move; the readout names the signature. The point is that each incident has a characteristic pattern across independent signals, and that pattern is what points at the fix. The long-prompt case below is the serving-side twin of long-context serving, where the KV cache rather than the queue is the limiting resource.
Per-request Gantt for the current run: grey queued, light prefill, dark decode, gold preempted. Inject a failure and watch the bands change.
| Signature | What moves | First check |
|---|---|---|
| Overload / queue saturation | Queue depth ↑, TTFT p99 ↑↑, goodput ↓, throughput flat | Admission control and autoscaling headroom |
| Capacity loss | Preemptions ↑, TPOT p99 ↑, KV occupancy ↑, batch occupancy ↓ | Replica health, then swap-vs-recompute policy |
| Long-prompt head-of-line | TTFT p50 ↑↑, ITL p99 spikes, short requests starve | Chunked prefill budget, then queueing policy |
| Prefix-cache thrash | Hit rate → 0, prefill volume ↑, KV pressure ↑, TTFT ↑ | Routing affinity and cache uniqueness |
| Retry storm | Error rate ↑, arrivals ↑, everything above at once | Client timeouts and backoff, not the model |
Routing, cascades, semantic caching and the batch tier
Spend where the difficulty is
A router sends easy requests to a cheap model and hard ones to an expensive one; a cascade sends everything cheap first and escalates when a confidence check fails. Both trace an accuracy-versus-cost frontier: as the routing threshold moves, you buy accuracy with money along a curve whose shape is set by the difficulty distribution. The classic failure is a cascade that escalates on a noisy signal — you pay both models on the same request, and cost rises while accuracy stalls.
Semantic caching answers a request from a near-duplicate already served, at near-zero model cost. It shifts the whole frontier toward lower cost, but it can also cap accuracy: a cached answer is only as good as the match. Cache the stable, high-frequency, low-risk requests, and make the similarity threshold a tunable so you can trade cost against a measured accuracy drop.
The batch tier is the last routing decision: non-interactive work — offline scoring, evaluation, embeddings, bulk summarisation — belongs on a queue with no latency SLO, where the scheduler can use every token of batch capacity. Moving even a modest share of traffic off the interactive pool is often the cheapest capacity you can buy, and it is invisible in a token price. The scheduler part owns the queueing mechanics; the frontier below owns the trade.
Accuracy vs cost. The faded curve is no cache; the solid curve is the current semantic-cache hit rate. The dot is the current routing threshold.
RPS → GPUs → dollars, and cost per task
The closing planner
Capacity planning is a chain: a workload in requests per second and token lengths becomes GPU-seconds per second (prefill and decode separately, because one is compute-bound and one is memory-bound), divided by a utilisation target; GPU count times an hourly price becomes cost per million tokens; and that becomes cost per task only after you know how many tokens a task consumes and how often it succeeds. Per-token prices hide the last step, which is the one users feel. A task's token count is itself variable: reasoning models emit long thinking traces before any answer, and agents spend several turns over a shared context. Pricing by the token averages those away.
Every lever the series introduced is a toggle here. GQA and MLA change the KV footprint (the mechanism is in the training guide's architecture chapter); KV and weight quantization shrink the bytes read per step; prefix caching removes prefill work; chunked prefill and disaggregation change how the two phases share a pool; speculation buys decode tokens; MoE serves only the active parameters. Each lever is annotated with the GPUs it is currently saving by itself, and the ranked list below the planner sorts them for the current inputs, so the biggest lever changes when you change the workload. The model is a first-order planner, not a substitute for a measured run — but the ranking is what you would spend the next benchmark on.
GPUs saved by each lever in isolation, at the current inputs, sorted. Longer bar = bigger lever.
Cheat sheet
| Concern | Rule |
|---|---|
| Open vs closed loop | Open loop: rate in, latency out. Closed loop: concurrency in, rate out, and the queue never builds. Quote the harness, concurrency and think time. |
| Harness choice | vllm bench for engine regressions · GenAI-Perf for endpoint concurrency · MLPerf Inference for comparable public claims. |
| Workload realism | Synthetic prompts understate prefill cost and overstate cache hit rate. Replay a sampled trace for capacity. |
| Warmup | Discard tens of seconds or a few hundred requests; measure a steady-state window; report both. |
| Throughput claim | Report token distributions, arrival process, SLO, stack/version and window. Anything less is not comparable. |
| Determinism | FP32 addition is not associative; batch shape changes reduction order and can flip an argmax at temperature 0. Batch-invariant kernels fix it. |
| Eight signals | Queue wait · TTFT p50/p99 · TPOT p50/p99 · goodput · KV occupancy/preemptions · prefix hit rate · batch occupancy/step time · tenant mix/retries. |
| Incident triage | Queue growth = overload · preemptions+TPOT = capacity loss · TTFT p50 + ITL spike = long prompt · hit rate → 0 = cache thrash · errors+arrivals = retry storm. |
| Cost | Cost per million tokens falls with throughput and utilisation; cost per successful task also moves with turns and retries. Optimise the task, not the token. |
Further reading
- Reddi et al., "MLPerf Inference Benchmark" (ISCA 2020) and the MLCommons Inference rules — the audited, comparable end of benchmarking.
- NVIDIA, GenAI-Perf — production endpoint load generation and per-request metrics.
- vLLM, benchmarking documentation for
vllm bench serve/throughput/latency. - Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference" (2025) — batch non-invariance and batch-invariant kernels; see also the vLLM batch-invariance discussion.
- Zhong et al., "DistServe: Disaggregating Prefill and Decoding for Goodput-optimized LLM Serving" (OSDI 2024) — the 7.4×/12.6× goodput result used by the disaggregation lever.
- Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention" (SOSP 2023) — the engine that made these measurements routine.
Cheat sheet
| Instrument | What it captures | Rule |
|---|---|---|
| Open-loop harness | Fixed arrival rate; latency is the output | Use for overload and tail claims |
| Closed-loop harness | Fixed clients; rate = clients ÷ (think + latency) | Optimistic p99; the plateau is the harness ceiling |
| Warmup | CUDA context, graph capture, autotuning | Discard it, and report the window you measured |
| Measurement card | Token distributions, arrival process, SLO, stack, window | Without all five, a number isn’t comparable |
| Batch invariance | Float addition isn’t associative | Greedy decoding isn’t reproducible across batch shapes |
| Capacity plan | RPS → GPU-seconds per request → GPUs → USD per task | Plan from a replayed trace, not averages |