Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The four numbers

One request, four clocks

TTFT (time to first token) is arrival-to-first-token: queueing, scheduling, the whole prompt prefill, and the first decode step. TPOT (time per output token), also called inter-token latency or ITL, is the steady-state gap between tokens after the first. Throughput is tokens per second produced by the whole deployment. Goodput is throughput that also meets the latency SLO — the only one of the four that a user can feel.

A single request is a sequence of phases, and the two headline metrics are measured at different places inside it. Drag the phase sliders to watch the measured TTFT and TPOT markers move; toggle streaming to see what "latency" means to a user watching tokens arrive. The training guide's evaluation chapter covers how a single benchmark number is constructed — serving metrics are the production analogue, sampled continuously rather than measured once.

One request’s timeline: queue, schedule, prefill, first token, decode steps, detokenize, flush. Marker lines are measured TTFT and TPOT.

2

Percentiles, not means

The tail is the product

A mean latency is the average of an experience almost nobody has. Latency distributions are right-skewed: a tight median with a long tail of queueing, preemption and retries. Percentiles describe the experience users actually report — p50 is the typical request, p99 is the one that generates the support ticket. This is why SLOs are written at p95 or p99, never at the mean.

The demo injects a 2% straggler population into a deterministic sample of 600 requests. The mean barely notices; the p99 moves several-fold. Every tail-hiding bug in a serving stack hides here, and the frontier in the next section is the tool for finding it.

Same sample, before and after 2% of requests become stragglers. Bars share a scale, so the p99 jump is not an artefact of normalisation.

💡 Rule of thumb: if your dashboard shows average latency, you are not measuring latency. Report p50, p95 and p99 together, and alert on the p99 burn rate rather than a mean threshold.
3

The throughput–latency frontier

The demo the series refers back to

Serving has no single "speed". Push more requests per second through a fixed deployment and throughput rises until the hardware saturates, then latency rises while throughput stops moving. The result is a curve, not a number: for every operating point there is a throughput and a latency distribution, and the useful question is where on that curve a deployment should sit for a given SLO.

This sweep runs the shared discrete-event simulator at twenty arrival rates — a Poisson workload, continuous batching, a real scheduler and KV capacity — and plots p50/p99 TTFT and TPOT against achieved tokens/s. Drag the two SLO lines (or the sliders) and the goodput card recomputes against the sampled requests. The knee is the point worth building a capacity plan around.

Log–log axes. Four latency series against achieved throughput; two draggable SLO thresholds; a marker for the selected operating load.

4

Goodput

Throughput that counts

Raw throughput rewards a server for generating tokens that arrive too late to be useful. Goodput applies the SLO as a filter first, then counts: the fraction of requests whose TTFT and TPOT both meet target, the requests/s that satisfy it, and the tokens/s they carry. A deployment can show 1,200 tokens/s of throughput and 40% goodput simultaneously — the other 60% are requests the user abandoned or the client timed out on.

Goodput is the objective a scheduler actually optimises, and it is the quantity that jumps when prefill and decode are split into separate pools. Watch it change in the frontier demo as the SLO lines move: tightening TPOT by 10 ms can cost a larger fraction of goodput than doubling capacity buys back.

💡 Define the unit: goodput is always "requests meeting both a TTFT and a TPOT target". A single latency target lets a system pass by being fast on the metric it controls and slow on the one the user feels.
5

What streaming does to "latency"

Perceived time is not total time

Without streaming, a user waits for the entire response before seeing anything, so perceived latency equals end-to-end time — often tens of seconds. With streaming, first content arrives at TTFT and the response grows under the user's eyes; perceived latency collapses to TTFT, and TPOT becomes the visible "typing speed". The same backend, the same total time, two completely different experiences.

This is why TTFT and TPOT are reported separately rather than collapsed into a single end-to-end number, and why a chunked prefill that spikes ITL is a visible regression even when throughput is flat. The streaming toggle in the first demo moves the "first content" marker; everything to its right is perceived as already-arriving work.

6

Where tail latency comes from

A checklist, not a mystery

Tail latency is rarely one bug; it is a queueing system telling you which resource ran out. The usual suspects: queueing when arrival rate approaches service rate (the frontier is steep here); head-of-line blocking when a long prompt or a long generation occupies the batch; preemption when KV pressure forces a running request to be recomputed or swapped; chunked-prefill ITL spikes when a large prompt is interleaved with decode; cold start when an autoscaler adds a replica and the first request pays for weight loading; and retries that multiply load exactly when the system is already saturated.

The scheduler decides which of these you get. p99 TTFT that degrades while p50 is flat almost always means the queue, not the kernels — check the frontier before optimising the model.

7

Writing a defensible SLO

Pick a percentile, a target and a window

An SLO is three numbers: a percentile, a threshold, and a measurement window. "p99 TTFT under 2 s over 28 days" is defensible; "fast" is not. Pair every latency target with a goodput floor so the system cannot pass by refusing work, and state the workload the number was measured on — prompt length, output length and arrival rate change it completely.

MetricTypical targetWhat it protects
TTFT p50 / p99200 ms / 2 sPerceived responsiveness, prefix-cache hits
TPOT p50 / p9920 ms / 80 msReading speed; streaming smoothness under load
Goodput fraction≥ 99% of admitted requestsThroughput that is actually useful
Error / timeout rate< 0.1%Retry storms and silent shedding

Numbers are illustrative; the structure is not. The glossary collects the definitions, and Part 20 turns them into capacity and cost.

8

tok/s/GPU and dollars per million tokens

The two numbers finance asks for

Throughput per GPU is the deliverable of every efficiency technique in this series. Cost per million output tokens is that number divided into the GPU's hourly price and its utilisation. A deployment at 1,000 tokens/s per GPU on a 2 USD/hr instance with 70% utilisation costs about 0.79 USD per million output tokens — a figure you can compare directly to an API list price. Move the sliders and watch the comparison change; the point is that cost is a consequence of goodput, not a separate problem.

Self-hosted cost per million output tokens against historical API list prices. Bars are USD per million output tokens.

Cheat sheet

MetricDefinitionWhy it matters
TTFTArrival → first token: queue, schedule, prefill, first stepDominated by queueing under load, not the kernels
TPOT / ITLSteady-state gap between tokensChunked-prefill spikes show up here
ThroughputTokens/s for the whole deploymentRewards tokens that arrive too late
GoodputThroughput from requests meeting both TTFT and TPOT targetsThe objective a scheduler actually optimises
Percentilesp50 typical, p99 the support ticketNever alert on the mean
SLOPercentile + threshold + window, with a goodput floorState the workload it was measured on
Cost per M tokensGPU USD/hr ÷ (tok/s × 3,600 × utilisation)Compare against API list prices
?

Check your understanding

0/5 answered