Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Workload shape is a serving parameter

RPS is the least useful number on the slide

Part 3 defined the four output metrics — TTFT, TPOT, end-to-end latency and goodput — and the training guide's inference page defined prefill versus decode. What that recap does not give you is the demand side. A request is not a unit: it is a prompt of some length, an output of some length, an arrival pattern, and a share of a shared prefix with everything around it. Two workloads can agree on 50 requests per second and disagree by an order of magnitude on every resource the fleet cares about. Decode cost scales with output tokens, prefill cost with input tokens, KV residency with input tokens times how long the request lives, and cache value with how much prefix is reused. Shape turns RPS into those four quantities, and shape is what this page models.

$$\text{GPUs}\;\propto\;\text{rate}\times\Big(\underbrace{\frac{E[\text{prompt}]}{v_{\text{prefill}}}}_{\text{prefill seconds}}+\underbrace{\frac{E[\text{output}]}{v_{\text{decode}}}}_{\text{decode seconds}}\Big)$$

The rest of this page is five shapes that each break that arithmetic in a different place: reasoning moves E[output], agents multiply turns and add idle KV, vision moves E[prompt] by thousands of tokens, and voice fixes the latency budget so tightly that only TTFT and the first audio chunk fit inside it.

2

Reasoning blows up the output-length distribution

Same RPS, eight times the GPUs

Chain-of-thought and test-time compute scale a model's quality by letting it generate more tokens before it answers; the training guide's reasoning page covers why the technique works. The serving consequence is blunt: the output distribution stops being a short chat response and becomes a long, heavy-tailed one. A chat turn of 200 tokens and a reasoning turn of 2,000 tokens are the same request count and the same concurrency cost at admission, but the reasoning turn holds a decode slot — and a KV block — ten times as long. Because decode throughput is memory-bandwidth-bound and shared, that extra residency directly multiplies the fleet.

The demo below runs the same discrete-event engine on three output distributions at the same arrival rate, and sizes the fleet from the simulated per-request service time. Nothing changes but the shape.

GPUs needed at the slider's RPS, from simulated service time at 70% target utilisation.

💡 The number to remember: at the same RPS, chat→reasoning is roughly an 8× fleet multiplier. Capacity planning that starts from requests per second and a single "average" output length will under-provision reasoning by almost an order of magnitude.
3

Thinking-token budgets

Saturating quality, linear cost, one knee

If reasoning helps, why not let every request think for 50,000 tokens? Because quality saturates while cost does not. On most verifiable tasks the accuracy curve is concave in thinking budget: the first few hundred tokens buy most of the gain, and beyond a knee each further block of tokens buys nearly nothing. Cost, meanwhile, is exactly linear in tokens generated: it is decode time, and decode time is GPU-seconds. The economically correct budget is not the model's maximum but the point where the marginal quality per token collapses — and that point is different for a math problem than for a support reply. A budget is therefore a per-route policy, and it belongs in the gateway next to the rate limit.

The demo models quality as a saturating exponential and adds a rework penalty for wrong answers, so the effective cost per task has a minimum — that minimum is the knee you would actually configure.

Quality (blue) and effective cost per task (pink) against the thinking budget.

4

Agents: many turns, long gaps, enormous shared prefixes

The request is not the unit of work — the session is

An agentic session is a loop: send a context, get a tool call, execute the tool for a while, append the result, repeat. That changes three quantities at once. First, a single user action becomes many model turns, so the gateway's RPM budget and the engine's admission both see a burst. Second, the turns are separated by tool gaps of hundreds of milliseconds to many seconds, during which the sequence is alive but idle — its KV is resident or evicted, but it is not decoding. Third, every turn resends a prefix that grows monotonically: system prompt, tool schemas, prior turns, tool outputs. The shared prefix is the same bytes on every hop, which is exactly what prefix caching was built for; the training guide's tools-and-agents page covers the loop itself. On realistic agent traces, KV cache hit rate is the difference between a system that works and one that thrashes — vLLM with Mooncake reported P50 TTFT falling 46× on agentic traces as hit rate rose from 1.7% to 92.2% (Part 8).

So the agent workload's defining resource is not FLOPs but KV residency across gaps. Hold it, and you pay memory for idle time; evict it, and you pay recompute on resume; offload it, and you pay transfer. Which is correct is a function of gap length, and that is the next demo.

5

KV residency across tool calls

Hold in HBM, evict and recompute, or offload

A 70B-class model with GQA carries about 0.33 MB of KV per token, so a 16K-token agent context is a little over 5 GB of resident state — real HBM that another session could be decoding into. Holding it across a 2.5 s tool gap costs about 13 GB·s of occupancy and buys a zero-cost resume. Evicting frees the state but makes the next turn re-prefill 16K tokens at maybe 8,000 tokens/s, a 2 s stall added to that turn's TTFT. Offloading to host DRAM frees HBM and avoids recompute, but pays a copy out and back over PCIe at ~25 GB/s, about 0.4 s round trip. The crossover is short: hold is clearly right for sub-second gaps and wrong for gaps of many seconds under memory pressure, which is why production systems tier KV rather than pick one policy globally.

One agent turn: prefill, decode, tool gap, repeat. Policy buttons change the gap.

Resource-seconds per turn against gap length; the crossover is where hold stops winning.

6

Image tokens and the prefill:decode inversion

A picture is thousands of prompt tokens

Vision enters the transformer as tokens. A ViT-style encoder patchifies the image — commonly 14-pixel patches pooled to a 28-pixel effective token — so a 1024×1024 image is on the order of 1,400 visual tokens before any tiling, and a zoomed crop scheme can push that past 5,000. Those tokens join the prompt: they are prefill, they are compute-bound on the way in, and they are KV resident for the rest of the request. The output, meanwhile, is often short — a caption, a Yes/No answer, a bounding box. That inverts the usual ratio. Text chat might be one prefill token per eight decode tokens; a high-resolution image question can be thirty prefill tokens per decode token, and the request's cost profile flips from decode-heavy to prefill-heavy. How a vision transformer turns pixels into those tokens at all, and how patch size, tiling and token merging set the budget, is built from zero in Multimodal Models, Part 6 and its predecessor on the ViT.

Prefill (blue) versus decode (pink) work for a text baseline and the current image.

The pool-sizing consequence is direct and is the subject of Part 13: if prefill work dominates, disaggregating a fixed ratio of prefill to decode workers is wrong — the vision traffic wants a much larger prefill pool, and because the encoder itself is a separate compute stage, it can be its own pool entirely.

7

Encoder disaggregation and caching

The image is often shared; the question is not

The vision encoder — patch embedding, the ViT stack, the projector into the language model's embedding space — is a self-contained compute stage that produces a fixed set of image tokens for a given image. That has two useful properties. It can be disaggregated: run encoders on their own pool, sized for image throughput, and pass the projected embeddings to the language prefill pool so the LLM GPUs never hold encoder weights. And it can be cached: the same image appearing in many requests (a product photo, a dashboard screenshot, a document page) encodes identically, so the embedding is a cache key. A vision cache keyed on the image hash collapses the encoder cost of repeated images to a lookup, and the language-side prefix cache then reuses the projected tokens for free. The two caches compose, and on image-heavy traffic the encoder is the stage where caching pays first because the encoder has no decode phase at all. For the from-scratch version of that stage — what the encoder computes, why its output is a cache key, and how the cost model is built — see Multimodal Models, Part 16.

💡 Rule of thumb: cache the image embedding, not the pixels. Two requests that hash to the same image but ask different questions share the entire encoder output and the entire image-token prefix; only the question differs.
8

Realtime audio is full-duplex and budget-bound

Everything must fit in 200 milliseconds

Realtime speech is the tightest latency shape in the catalogue. A conversation feels natural when the gap between the user finishing and the assistant starting is under roughly 200 ms; past that, turn-taking breaks down. Inside that budget sits automatic speech recognition streaming a partial transcript, the network hop to the model, TTFT for the reply, and the text-to-speech first audio chunk. The budget is per direction and the system is full-duplex — the assistant may be speaking while it listens — so ASR and TTS run concurrently and share the GPU with decode. This is the shape that most strongly punishes a monolithic pool: a long text generation queued alongside a voice session will blow the 200 ms budget even if the average load is low, because the voice session's deadline is not a percentile target, it is a hard interaction constraint. The models inside that budget — CTC, RNN-T, Whisper, and the two-pass fix that makes them stream — are built from scratch in Generative Media, Part 19.

The 200 ms budget, filled with ASR, network, TTFT and TTS first chunk.

9

What each shape does to pool sizing

One fleet, four different bottlenecks

Put the shapes side by side and the pool-sizing answer stops being a single number. Reasoning is decode-bound: the prefill pool is small, the decode pool is large, and the binding constraint is decode batch slots and KV bandwidth. Agents are residency-bound: decode throughput matters, but the pool must also be sized for the KV of idle sessions and for the prefill spikes when caches miss, and the right architecture is prefix-cache-aware routing plus a KV tier. Vision is prefill-bound and encoder-heavy: it wants a large prefill pool, a separate encoder pool, and image-embedding cache in front of both. Voice is latency-bound: it wants a small dedicated pool with strict admission, because mixing it into a shared decode pool trades away the one guarantee its users can feel. The same accelerator, four different shapes, four different ratios — and the disaggregation ratio from Part 13 is the parameter you retune per route, not per cluster.

⚠ Sizing rule: split traffic into these four classes before you size anything. A blended average across them is the single most common capacity-planning error in serving, and it fails in exactly the direction of the loudest workload.
✓

Cheat sheet

Shape to bottleneck to pool

ShapeWhat movesBinding resourcePool response
ChatShort output, short promptDecode slots at low TPOTBalanced pools; the baseline
CodeMedium output, tighter TPOT SLODecode bandwidth; speculative valueLarge decode pool; speculation on
ReasoningOutput mean ×8, heavy tailDecode slots and KV residency~8× decode fleet; hard token budgets
AgentsMany turns; long idle gaps; huge shared prefixKV residency + prefix hit rateCache-aware routing; KV tiering; prefill headroom
VisionPrompt ×103 tokens per imagePrefill and encoder computeLarge prefill pool; separate encoder pool; embedding cache
Realtime voiceFixed 200 ms end-to-end budgetLatency, full-duplex GPU sharingSmall dedicated pool; strict admission
📚

Further reading

References

?

Check your understanding

0/5 answered