Reasoning, agents and multimodal
Two systems with the same model and the same RPS can need eight times different GPU fleets, because the workload shape — how many tokens come in, how many go out, how often, and whether an image arrives with them — is a serving parameter in its own right. This part makes shape measurable: output-length distributions, thinking budgets, agent turns with tool gaps, image tokens, and the hard 200 ms wall of realtime voice.
Workload shape is a serving parameter
RPS is the least useful number on the slide
Part 3 defined the four output metrics — TTFT, TPOT, end-to-end latency and goodput — and the training guide's inference page defined prefill versus decode. What that recap does not give you is the demand side. A request is not a unit: it is a prompt of some length, an output of some length, an arrival pattern, and a share of a shared prefix with everything around it. Two workloads can agree on 50 requests per second and disagree by an order of magnitude on every resource the fleet cares about. Decode cost scales with output tokens, prefill cost with input tokens, KV residency with input tokens times how long the request lives, and cache value with how much prefix is reused. Shape turns RPS into those four quantities, and shape is what this page models.
The rest of this page is five shapes that each break that arithmetic in a different place: reasoning moves E[output], agents multiply turns and add idle KV, vision moves E[prompt] by thousands of tokens, and voice fixes the latency budget so tightly that only TTFT and the first audio chunk fit inside it.
Reasoning blows up the output-length distribution
Same RPS, eight times the GPUs
Chain-of-thought and test-time compute scale a model's quality by letting it generate more tokens before it answers; the training guide's reasoning page covers why the technique works. The serving consequence is blunt: the output distribution stops being a short chat response and becomes a long, heavy-tailed one. A chat turn of 200 tokens and a reasoning turn of 2,000 tokens are the same request count and the same concurrency cost at admission, but the reasoning turn holds a decode slot — and a KV block — ten times as long. Because decode throughput is memory-bandwidth-bound and shared, that extra residency directly multiplies the fleet.
The demo below runs the same discrete-event engine on three output distributions at the same arrival rate, and sizes the fleet from the simulated per-request service time. Nothing changes but the shape.
GPUs needed at the slider's RPS, from simulated service time at 70% target utilisation.
Thinking-token budgets
Saturating quality, linear cost, one knee
If reasoning helps, why not let every request think for 50,000 tokens? Because quality saturates while cost does not. On most verifiable tasks the accuracy curve is concave in thinking budget: the first few hundred tokens buy most of the gain, and beyond a knee each further block of tokens buys nearly nothing. Cost, meanwhile, is exactly linear in tokens generated: it is decode time, and decode time is GPU-seconds. The economically correct budget is not the model's maximum but the point where the marginal quality per token collapses — and that point is different for a math problem than for a support reply. A budget is therefore a per-route policy, and it belongs in the gateway next to the rate limit.
The demo models quality as a saturating exponential and adds a rework penalty for wrong answers, so the effective cost per task has a minimum — that minimum is the knee you would actually configure.
Quality (blue) and effective cost per task (pink) against the thinking budget.
Agents: many turns, long gaps, enormous shared prefixes
The request is not the unit of work — the session is
An agentic session is a loop: send a context, get a tool call, execute the tool for a while, append the result, repeat. That changes three quantities at once. First, a single user action becomes many model turns, so the gateway's RPM budget and the engine's admission both see a burst. Second, the turns are separated by tool gaps of hundreds of milliseconds to many seconds, during which the sequence is alive but idle — its KV is resident or evicted, but it is not decoding. Third, every turn resends a prefix that grows monotonically: system prompt, tool schemas, prior turns, tool outputs. The shared prefix is the same bytes on every hop, which is exactly what prefix caching was built for; the training guide's tools-and-agents page covers the loop itself. On realistic agent traces, KV cache hit rate is the difference between a system that works and one that thrashes — vLLM with Mooncake reported P50 TTFT falling 46× on agentic traces as hit rate rose from 1.7% to 92.2% (Part 8).
So the agent workload's defining resource is not FLOPs but KV residency across gaps. Hold it, and you pay memory for idle time; evict it, and you pay recompute on resume; offload it, and you pay transfer. Which is correct is a function of gap length, and that is the next demo.
KV residency across tool calls
Hold in HBM, evict and recompute, or offload
A 70B-class model with GQA carries about 0.33 MB of KV per token, so a 16K-token agent context is a little over 5 GB of resident state — real HBM that another session could be decoding into. Holding it across a 2.5 s tool gap costs about 13 GB·s of occupancy and buys a zero-cost resume. Evicting frees the state but makes the next turn re-prefill 16K tokens at maybe 8,000 tokens/s, a 2 s stall added to that turn's TTFT. Offloading to host DRAM frees HBM and avoids recompute, but pays a copy out and back over PCIe at ~25 GB/s, about 0.4 s round trip. The crossover is short: hold is clearly right for sub-second gaps and wrong for gaps of many seconds under memory pressure, which is why production systems tier KV rather than pick one policy globally.
One agent turn: prefill, decode, tool gap, repeat. Policy buttons change the gap.
Resource-seconds per turn against gap length; the crossover is where hold stops winning.
Image tokens and the prefill:decode inversion
A picture is thousands of prompt tokens
Vision enters the transformer as tokens. A ViT-style encoder patchifies the image — commonly 14-pixel patches pooled to a 28-pixel effective token — so a 1024×1024 image is on the order of 1,400 visual tokens before any tiling, and a zoomed crop scheme can push that past 5,000. Those tokens join the prompt: they are prefill, they are compute-bound on the way in, and they are KV resident for the rest of the request. The output, meanwhile, is often short — a caption, a Yes/No answer, a bounding box. That inverts the usual ratio. Text chat might be one prefill token per eight decode tokens; a high-resolution image question can be thirty prefill tokens per decode token, and the request's cost profile flips from decode-heavy to prefill-heavy. How a vision transformer turns pixels into those tokens at all, and how patch size, tiling and token merging set the budget, is built from zero in Multimodal Models, Part 6 and its predecessor on the ViT.
Prefill (blue) versus decode (pink) work for a text baseline and the current image.
The pool-sizing consequence is direct and is the subject of Part 13: if prefill work dominates, disaggregating a fixed ratio of prefill to decode workers is wrong — the vision traffic wants a much larger prefill pool, and because the encoder itself is a separate compute stage, it can be its own pool entirely.
Encoder disaggregation and caching
The image is often shared; the question is not
The vision encoder — patch embedding, the ViT stack, the projector into the language model's embedding space — is a self-contained compute stage that produces a fixed set of image tokens for a given image. That has two useful properties. It can be disaggregated: run encoders on their own pool, sized for image throughput, and pass the projected embeddings to the language prefill pool so the LLM GPUs never hold encoder weights. And it can be cached: the same image appearing in many requests (a product photo, a dashboard screenshot, a document page) encodes identically, so the embedding is a cache key. A vision cache keyed on the image hash collapses the encoder cost of repeated images to a lookup, and the language-side prefix cache then reuses the projected tokens for free. The two caches compose, and on image-heavy traffic the encoder is the stage where caching pays first because the encoder has no decode phase at all. For the from-scratch version of that stage — what the encoder computes, why its output is a cache key, and how the cost model is built — see Multimodal Models, Part 16.
Realtime audio is full-duplex and budget-bound
Everything must fit in 200 milliseconds
Realtime speech is the tightest latency shape in the catalogue. A conversation feels natural when the gap between the user finishing and the assistant starting is under roughly 200 ms; past that, turn-taking breaks down. Inside that budget sits automatic speech recognition streaming a partial transcript, the network hop to the model, TTFT for the reply, and the text-to-speech first audio chunk. The budget is per direction and the system is full-duplex — the assistant may be speaking while it listens — so ASR and TTS run concurrently and share the GPU with decode. This is the shape that most strongly punishes a monolithic pool: a long text generation queued alongside a voice session will blow the 200 ms budget even if the average load is low, because the voice session's deadline is not a percentile target, it is a hard interaction constraint. The models inside that budget — CTC, RNN-T, Whisper, and the two-pass fix that makes them stream — are built from scratch in Generative Media, Part 19.
The 200 ms budget, filled with ASR, network, TTFT and TTS first chunk.
What each shape does to pool sizing
One fleet, four different bottlenecks
Put the shapes side by side and the pool-sizing answer stops being a single number. Reasoning is decode-bound: the prefill pool is small, the decode pool is large, and the binding constraint is decode batch slots and KV bandwidth. Agents are residency-bound: decode throughput matters, but the pool must also be sized for the KV of idle sessions and for the prefill spikes when caches miss, and the right architecture is prefix-cache-aware routing plus a KV tier. Vision is prefill-bound and encoder-heavy: it wants a large prefill pool, a separate encoder pool, and image-embedding cache in front of both. Voice is latency-bound: it wants a small dedicated pool with strict admission, because mixing it into a shared decode pool trades away the one guarantee its users can feel. The same accelerator, four different shapes, four different ratios — and the disaggregation ratio from Part 13 is the parameter you retune per route, not per cluster.
Cheat sheet
Shape to bottleneck to pool
| Shape | What moves | Binding resource | Pool response |
|---|---|---|---|
| Chat | Short output, short prompt | Decode slots at low TPOT | Balanced pools; the baseline |
| Code | Medium output, tighter TPOT SLO | Decode bandwidth; speculative value | Large decode pool; speculation on |
| Reasoning | Output mean ×8, heavy tail | Decode slots and KV residency | ~8× decode fleet; hard token budgets |
| Agents | Many turns; long idle gaps; huge shared prefix | KV residency + prefix hit rate | Cache-aware routing; KV tiering; prefill headroom |
| Vision | Prompt ×103 tokens per image | Prefill and encoder compute | Large prefill pool; separate encoder pool; embedding cache |
| Realtime voice | Fixed 200 ms end-to-end budget | Latency, full-duplex GPU sharing | Small dedicated pool; strict admission |
Further reading
References
- Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, NeurIPS 2022.
- Pope et al., "Efficiently Scaling Transformer Inference" (MLSys 2023) — the service-time model behind the fleet arithmetic.
- Dosovitskiy et al., An Image is Worth 16x16 Words, ICLR 2021 — patchification and the visual-token count.
- vLLM x Mooncake: Serving Agentic Workloads at Scale — prefix-cache hit rate and TTFT on agentic traces.
- The disaggregation part for sizing prefill and decode pools, and prefix caching for the KV hierarchy agent traffic depends on.