Inference, Serving & Efficiency
Training produces a fixed set of weights. Everything from here is about the very different problem of running that fixed function fast, cheaply, and for many users at once — which turns out to have its own rich set of mechanisms, none of them touched by training itself.
Choosing the next token: the full sampler
Decoding
Part 1 sampled greedily; a real serving stack layers several filters onto the softmax distribution before sampling. Each rule below removes or reshapes probability mass differently — toggle them to see exactly which tokens survive.
| Rule | What it removes |
|---|---|
| Temperature | Nothing removed — reshapes sharpness (÷T before softmax) |
| Top-k | Everything outside the k highest-probability tokens |
| Top-p / nucleus | Tokens outside the smallest set whose cumulative probability ≥ p |
| Min-p | Tokens whose probability is below a fraction of the top token's probability — adapts to how peaked the distribution is, unlike a fixed top-k |
| Repetition penalty | Down-weights tokens already generated earlier in the response |
The KV cache
Making autoregressive decoding tractable
Generating token t+1 needs every earlier token's key and value vectors for attention. Recomputing them from scratch at every new token would make generation quadratic in sequence length for no reason — so instead, every key/value computed is cached and reused, and only the newest token's K/V is computed each step. The cost: the cache itself takes real memory, growing linearly with context length, layers, and KV heads.
Prefill vs. decode
Two very different workloads in one request
Prefill processes the whole prompt at once — one matrix multiply over many tokens, which keeps the GPU's compute units busy (compute-bound; high arithmetic intensity). Decode generates one token at a time — each step reloads the full model weights (and KV cache) from GPU memory to produce a single new token, so it's bottlenecked on memory bandwidth, not compute (memory-bound; low arithmetic intensity). This is why decode throughput barely improves on a faster-but-not-more-bandwidth GPU, and why batching many concurrent requests' decode steps together (below) is the main lever for decode efficiency.
Continuous batching
Serving many requests at once
Static batching waits for a fixed group of requests to all finish before starting the next batch — one long response blocks the whole batch's GPU slot from being reused. Continuous batching (also called in-flight batching) instead swaps a finished request out and a new one in at every decode step, keeping the GPU's batch dimension full. Click below to compare.
Quantization
Trading precision for memory and speed
Weights trained in bf16 can be stored (and often computed) in fewer bits at serving time — int8 or int4 — shrinking memory and often improving throughput, at some cost in output quality. Below, a synthetic weight distribution (roughly what a trained layer's weights look like) is quantized at different bit widths; watch the size shrink and the reconstruction error (mean squared error against the original) grow.
Real quantization methods (GPTQ, AWQ) are smarter than uniform rounding — they choose scales per-channel or per-group and account for which weights matter most to the output — but the basic size/error trade-off shown here is the same one they're managing.
Further serving techniques
Named, not demoed
Serving at scale: inference for RL
Why the biggest serving customer is training
The largest recent consumer of inference is not a chatbot endpoint — it's reinforcement learning. Generating long chains of thought means rollouts up to 32K tokens (mean ≈ 14,628 for a reasoning model), and the learner spends roughly 75% of its time waiting for data. In OLMo 3's 32B run, 8 H100 nodes trained while 20 generated rollouts (5× the inference compute); the 7B run used 2 training and 7 inference nodes (14×). Every serving technique from this page therefore shows up directly inside the RL system — and in fact the fixes are the same continuous-batching and scheduling machinery, adapted to a workload where sequences are highly variable in length and the "requests" are policy samples.
OLMoRL training throughput as each serving/systems improvement is added. MBU = memory-bandwidth utilization; MFU = model FLOPs utilization.
| Change | Why it matters |
|---|---|
| Continuous batching | Rollouts have wildly different lengths. Static batching wastes ≈54% of slot-steps waiting for the longest one (mean 14,628 vs max 32,768 tokens); continuous batching backfills freed slots immediately. |
| Active sampling | Instead of DAPO's fixed 3× oversampling, pull completed prompt–completion pairs continuously from a result queue — stable batch size and loss even as filtering drops easy prompts. |
| Inflight weight updates | Push updated policy weights to the inference workers without pausing generation or invalidating the KV cache. Up to 4× faster than stopping to sync. |
Stacked together, these took training throughput from 881 to 2,949 tokens/second and memory-bandwidth utilization from 12.9% to 43.2%. The lesson generalizes: once long-context generation is in the loop, RL is an inference-systems problem first and an optimization problem second.
Cheat sheet
Recap
| Concept | What it is |
|---|---|
| KV cache | Cached per-token key/value vectors so decode doesn't recompute the whole sequence's attention every step |
| Prefill vs. decode | Compute-bound (whole prompt at once) vs. memory-bandwidth-bound (one token at a time) |
| Continuous batching | Swap finished requests out and new ones in every decode step, not per fixed batch |
| Quantization | Storing/computing weights at lower precision (int8/int4) to save memory and increase throughput |
| Speculative decoding | A small model drafts tokens; the large model verifies them in one pass |
Further reading
References
- Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention" (2023) — vLLM.
- Yu et al., "Orca: A Distributed Serving System for Transformer-Based Generative Models" (2022) — continuous batching.
- Leviathan et al., "Fast Inference from Transformers via Speculative Decoding" (2023).
- Frantar et al., "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers" (2023).