TTFT, TPOT and goodput
A serving system is described by four numbers, and none of them is "latency". This part defines them, shows why the mean is the wrong statistic, builds the throughput–latency frontier the rest of the series keeps referring back to, and ends with the two numbers finance actually asks for.
The four numbers
One request, four clocks
TTFT (time to first token) is arrival-to-first-token: queueing, scheduling, the whole prompt prefill, and the first decode step. TPOT (time per output token), also called inter-token latency or ITL, is the steady-state gap between tokens after the first. Throughput is tokens per second produced by the whole deployment. Goodput is throughput that also meets the latency SLO — the only one of the four that a user can feel.
A single request is a sequence of phases, and the two headline metrics are measured at different places inside it. Drag the phase sliders to watch the measured TTFT and TPOT markers move; toggle streaming to see what "latency" means to a user watching tokens arrive. The training guide's evaluation chapter covers how a single benchmark number is constructed — serving metrics are the production analogue, sampled continuously rather than measured once.
One request’s timeline: queue, schedule, prefill, first token, decode steps, detokenize, flush. Marker lines are measured TTFT and TPOT.
Percentiles, not means
The tail is the product
A mean latency is the average of an experience almost nobody has. Latency distributions are right-skewed: a tight median with a long tail of queueing, preemption and retries. Percentiles describe the experience users actually report — p50 is the typical request, p99 is the one that generates the support ticket. This is why SLOs are written at p95 or p99, never at the mean.
The demo injects a 2% straggler population into a deterministic sample of 600 requests. The mean barely notices; the p99 moves several-fold. Every tail-hiding bug in a serving stack hides here, and the frontier in the next section is the tool for finding it.
Same sample, before and after 2% of requests become stragglers. Bars share a scale, so the p99 jump is not an artefact of normalisation.
The throughput–latency frontier
The demo the series refers back to
Serving has no single "speed". Push more requests per second through a fixed deployment and throughput rises until the hardware saturates, then latency rises while throughput stops moving. The result is a curve, not a number: for every operating point there is a throughput and a latency distribution, and the useful question is where on that curve a deployment should sit for a given SLO.
This sweep runs the shared discrete-event simulator at twenty arrival rates — a Poisson workload, continuous batching, a real scheduler and KV capacity — and plots p50/p99 TTFT and TPOT against achieved tokens/s. Drag the two SLO lines (or the sliders) and the goodput card recomputes against the sampled requests. The knee is the point worth building a capacity plan around.
Log–log axes. Four latency series against achieved throughput; two draggable SLO thresholds; a marker for the selected operating load.
Goodput
Throughput that counts
Raw throughput rewards a server for generating tokens that arrive too late to be useful. Goodput applies the SLO as a filter first, then counts: the fraction of requests whose TTFT and TPOT both meet target, the requests/s that satisfy it, and the tokens/s they carry. A deployment can show 1,200 tokens/s of throughput and 40% goodput simultaneously — the other 60% are requests the user abandoned or the client timed out on.
Goodput is the objective a scheduler actually optimises, and it is the quantity that jumps when prefill and decode are split into separate pools. Watch it change in the frontier demo as the SLO lines move: tightening TPOT by 10 ms can cost a larger fraction of goodput than doubling capacity buys back.
What streaming does to "latency"
Perceived time is not total time
Without streaming, a user waits for the entire response before seeing anything, so perceived latency equals end-to-end time — often tens of seconds. With streaming, first content arrives at TTFT and the response grows under the user's eyes; perceived latency collapses to TTFT, and TPOT becomes the visible "typing speed". The same backend, the same total time, two completely different experiences.
This is why TTFT and TPOT are reported separately rather than collapsed into a single end-to-end number, and why a chunked prefill that spikes ITL is a visible regression even when throughput is flat. The streaming toggle in the first demo moves the "first content" marker; everything to its right is perceived as already-arriving work.
Where tail latency comes from
A checklist, not a mystery
Tail latency is rarely one bug; it is a queueing system telling you which resource ran out. The usual suspects: queueing when arrival rate approaches service rate (the frontier is steep here); head-of-line blocking when a long prompt or a long generation occupies the batch; preemption when KV pressure forces a running request to be recomputed or swapped; chunked-prefill ITL spikes when a large prompt is interleaved with decode; cold start when an autoscaler adds a replica and the first request pays for weight loading; and retries that multiply load exactly when the system is already saturated.
The scheduler decides which of these you get. p99 TTFT that degrades while p50 is flat almost always means the queue, not the kernels — check the frontier before optimising the model.
Writing a defensible SLO
Pick a percentile, a target and a window
An SLO is three numbers: a percentile, a threshold, and a measurement window. "p99 TTFT under 2 s over 28 days" is defensible; "fast" is not. Pair every latency target with a goodput floor so the system cannot pass by refusing work, and state the workload the number was measured on — prompt length, output length and arrival rate change it completely.
| Metric | Typical target | What it protects |
|---|---|---|
| TTFT p50 / p99 | 200 ms / 2 s | Perceived responsiveness, prefix-cache hits |
| TPOT p50 / p99 | 20 ms / 80 ms | Reading speed; streaming smoothness under load |
| Goodput fraction | ≥ 99% of admitted requests | Throughput that is actually useful |
| Error / timeout rate | < 0.1% | Retry storms and silent shedding |
Numbers are illustrative; the structure is not. The glossary collects the definitions, and Part 20 turns them into capacity and cost.
tok/s/GPU and dollars per million tokens
The two numbers finance asks for
Throughput per GPU is the deliverable of every efficiency technique in this series. Cost per million output tokens is that number divided into the GPU's hourly price and its utilisation. A deployment at 1,000 tokens/s per GPU on a 2 USD/hr instance with 70% utilisation costs about 0.79 USD per million output tokens — a figure you can compare directly to an API list price. Move the sliders and watch the comparison change; the point is that cost is a consequence of goodput, not a separate problem.
Self-hosted cost per million output tokens against historical API list prices. Bars are USD per million output tokens.
Cheat sheet
| Metric | Definition | Why it matters |
|---|---|---|
| TTFT | Arrival → first token: queue, schedule, prefill, first step | Dominated by queueing under load, not the kernels |
| TPOT / ITL | Steady-state gap between tokens | Chunked-prefill spikes show up here |
| Throughput | Tokens/s for the whole deployment | Rewards tokens that arrive too late |
| Goodput | Throughput from requests meeting both TTFT and TPOT targets | The objective a scheduler actually optimises |
| Percentiles | p50 typical, p99 the support ticket | Never alert on the mean |
| SLO | Percentile + threshold + window, with a goodput floor | State the workload it was measured on |
| Cost per M tokens | GPU USD/hr ÷ (tok/s × 3,600 × utilisation) | Compare against API list prices |