Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Choosing the next token: the full sampler

Decoding

Part 1 sampled greedily; a real serving stack layers several filters onto the softmax distribution before sampling. Each rule below removes or reshapes probability mass differently — toggle them to see exactly which tokens survive.

RuleWhat it removes
TemperatureNothing removed — reshapes sharpness (÷T before softmax)
Top-kEverything outside the k highest-probability tokens
Top-p / nucleusTokens outside the smallest set whose cumulative probability ≥ p
Min-pTokens whose probability is below a fraction of the top token's probability — adapts to how peaked the distribution is, unlike a fixed top-k
Repetition penaltyDown-weights tokens already generated earlier in the response
2

The KV cache

Making autoregressive decoding tractable

Generating token t+1 needs every earlier token's key and value vectors for attention. Recomputing them from scratch at every new token would make generation quadratic in sequence length for no reason — so instead, every key/value computed is cached and reused, and only the newest token's K/V is computed each step. The cost: the cache itself takes real memory, growing linearly with context length, layers, and KV heads.

$$\text{KV cache size} = 2 \times L \times n_{kv} \times d_{head} \times T \times \text{bytes} \times \text{batch}$$
💡 Why GQA (Part 3) matters here: dropping from 32 independent KV heads to, say, 4 shared groups shrinks the cache by 8× for free — this is the actual production motivation for grouped-query attention, not just a training-time nicety.
3

Prefill vs. decode

Two very different workloads in one request

Prefill processes the whole prompt at once — one matrix multiply over many tokens, which keeps the GPU's compute units busy (compute-bound; high arithmetic intensity). Decode generates one token at a time — each step reloads the full model weights (and KV cache) from GPU memory to produce a single new token, so it's bottlenecked on memory bandwidth, not compute (memory-bound; low arithmetic intensity). This is why decode throughput barely improves on a faster-but-not-more-bandwidth GPU, and why batching many concurrent requests' decode steps together (below) is the main lever for decode efficiency.

4

Continuous batching

Serving many requests at once

Static batching waits for a fixed group of requests to all finish before starting the next batch — one long response blocks the whole batch's GPU slot from being reused. Continuous batching (also called in-flight batching) instead swaps a finished request out and a new one in at every decode step, keeping the GPU's batch dimension full. Click below to compare.

5

Quantization

Trading precision for memory and speed

Weights trained in bf16 can be stored (and often computed) in fewer bits at serving time — int8 or int4 — shrinking memory and often improving throughput, at some cost in output quality. Below, a synthetic weight distribution (roughly what a trained layer's weights look like) is quantized at different bit widths; watch the size shrink and the reconstruction error (mean squared error against the original) grow.

Real quantization methods (GPTQ, AWQ) are smarter than uniform rounding — they choose scales per-channel or per-group and account for which weights matter most to the output — but the basic size/error trade-off shown here is the same one they're managing.

6

Further serving techniques

Named, not demoed

Speculative decoding — a small "draft" model proposes several tokens ahead; the large model verifies them in one batched forward pass, accepting the correct prefix. When the draft model's guesses are often right, this turns several slow decode steps into one, at the cost of running two models.
Context-length extension — techniques like position interpolation or YaRN rescale RoPE's rotation frequencies so a model trained at, say, 4K context can be used (often with brief fine-tuning) well beyond it, at some quality cost near the new limit.
Constrained / structured decoding — restricting the sampler to only tokens that keep a generation valid against a grammar or JSON schema, by masking out disallowed tokens at each step — how "guaranteed valid JSON" output modes work.
Tool calling — the model emits a structured request (usually constrained-decoded) that the serving stack intercepts, executes against a real function or API, and feeds the result back in as additional context before continuing generation.
RAG (retrieval-augmented generation) — before generation starts, a retrieval system (often the same embedding idea from Part 2, at a much larger scale) finds relevant documents and prepends them to the prompt, so the model can answer from up-to-date or private text it never saw in pretraining.
Beam search — instead of sampling one path, keep the top-B partial sequences by cumulative probability at every step; used less often for open-ended chat (where diverse, sampled text feels better) than for tasks with a single correct answer, like translation.
7

Serving at scale: inference for RL

Why the biggest serving customer is training

The largest recent consumer of inference is not a chatbot endpoint — it's reinforcement learning. Generating long chains of thought means rollouts up to 32K tokens (mean ≈ 14,628 for a reasoning model), and the learner spends roughly 75% of its time waiting for data. In OLMo 3's 32B run, 8 H100 nodes trained while 20 generated rollouts (5× the inference compute); the 7B run used 2 training and 7 inference nodes (14×). Every serving technique from this page therefore shows up directly inside the RL system — and in fact the fixes are the same continuous-batching and scheduling machinery, adapted to a workload where sequences are highly variable in length and the "requests" are policy samples.

OLMoRL training throughput as each serving/systems improvement is added. MBU = memory-bandwidth utilization; MFU = model FLOPs utilization.

ChangeWhy it matters
Continuous batchingRollouts have wildly different lengths. Static batching wastes ≈54% of slot-steps waiting for the longest one (mean 14,628 vs max 32,768 tokens); continuous batching backfills freed slots immediately.
Active samplingInstead of DAPO's fixed 3× oversampling, pull completed prompt–completion pairs continuously from a result queue — stable batch size and loss even as filtering drops easy prompts.
Inflight weight updatesPush updated policy weights to the inference workers without pausing generation or invalidating the KV cache. Up to 4× faster than stopping to sync.

Stacked together, these took training throughput from 881 to 2,949 tokens/second and memory-bandwidth utilization from 12.9% to 43.2%. The lesson generalizes: once long-context generation is in the loop, RL is an inference-systems problem first and an optimization problem second.

📌 Cross-link: Part 15 shows the full OLMoRL recipe and the batching timeline this section summarizes.
✓

Cheat sheet

Recap

ConceptWhat it is
KV cacheCached per-token key/value vectors so decode doesn't recompute the whole sequence's attention every step
Prefill vs. decodeCompute-bound (whole prompt at once) vs. memory-bandwidth-bound (one token at a time)
Continuous batchingSwap finished requests out and new ones in every decode step, not per fixed batch
QuantizationStoring/computing weights at lower precision (int8/int4) to save memory and increase throughput
Speculative decodingA small model drafts tokens; the large model verifies them in one pass
📚

Further reading

References

?

Check your understanding

0/5 answered
Speed and cost solved, the remaining question is how anyone knows whether the model is actually any good. The sequel — LLM Serving, Interactively — takes this part's sampler, KV cache and batching and rebuilds them as a quantitative model you plug your own hardware into, all the way out to disaggregated prefill/decode pools and $/M tokens. Continue: evaluation →