Continuous batching and chunked prefill
Part 5 made the KV cache allocatable. This part spends it. The question is what exactly the GPU computes in each iteration: a fixed group of requests that share a fate, or a fresh, ragged batch assembled at every step. The difference between those two answers is most of a serving engine's throughput — and the price is a scheduling problem in which a single long prompt can stall every decoder in the batch.
Static batching and the padding tax
One batch, one finish line
The oldest way to batch, and the one almost every other accelerator workload uses, is static batching: collect N requests, stack them, run the batch until every member has finished, then start the next batch. Matrix kernels need rectangular inputs, so the short sequences are padded out to the length of the longest one and the extra positions masked. Masking the arithmetic does not remove it: padded tokens cost the same FLOPs and the same attention reads as real ones. This is the padding tax, and it is largest exactly where chatbots live, in requests whose output lengths vary by an order of magnitude.
The second cost is worse than the padding. A batch cannot release a slot until its slowest member finishes, so a single 2,000-token response holds five finished 100-token responses hostage. The GPU's batch dimension — the thing that amortises a memory-bound decode step across requests — sits partly empty for the whole tail of every batch. Drag the variance slider below to see the tax itself; the Gantt in section 3 shows what it does to the timeline.
Six sequences padded to the longest in the batch. Hatched cells are masked padding that still cost compute.
Dynamic batching
Stop waiting for a full batch
Dynamic batching keeps the lockstep execution but changes when a batch is dispatched. Instead of waiting for exactly N requests, the scheduler starts a batch as soon as either N have queued or a maximum queue delay has elapsed, whichever comes first. At low load this dispatches small batches immediately, which is a large latency win; at high load it behaves like static batching, because N arrives long before the timer fires.
What it does not change is the finish line. The batch still runs to completion together, so padding and tail waste remain, and the GPU still idles between batches. Dynamic batching is the right tool for vision models with near-constant latency; for autoregressive decoding, where request durations differ by 100×, it is a half-measure. The demo below lets the same workload run under both regimes and under iteration-level batching.
Iteration-level batching (Orca)
The batch is rebuilt every token
Orca's observation, in 2022, was that a transformer decode step is already an iteration boundary: the model produces one token per sequence, and then the scheduler is free to change who is in the batch. Iteration-level batching (equivalently continuous or in-flight batching) does exactly that. Finished sequences leave the batch the moment they emit their stop token; waiting sequences join the moment the KV pool and the sequence cap allow. The batch dimension stays nearly full for the entire run, and no request waits for a stranger's last token.
The training guide's continuous batching demo shows the mechanism on a handful of requests. What it does not show is what the scheduler is optimising — that is this part's job — nor the two knobs that keep the iteration from becoming either idle or overstuffed, nor the prefill problem that arrives with it. The Gantt chart below runs one fixed request stream through all three policies.
What batch size does to step time
Flat, then linear, then you pay
Part 1's memory wall says a decode step reads the weights once per token no matter how large the batch. So adding the second sequence to a decode batch is almost free: the same 70 GB of weights streams from HBM while the tensor cores process two rows instead of one. This is why decode throughput scales nearly linearly with batch at small batch sizes. It stops when the batch's arithmetic intensity reaches the machine's ridge point — the batch size at which 2 × params × batch FLOPs take longer than streaming weights + KV bytes. Past that knee the step is compute-bound and its time grows linearly in batch.
The knee is a property of the machine and the model, not a tuning decision: for an 8B model at BF16 on H100 it lands near batch 295 with no cache, and the KV cache drags it down toward the batch sizes people actually use. The step time itself — the quantity that becomes the inter-token latency — is what the chart plots. Slide batch, and slide the cache per sequence, and watch the knee move.
Decode step time versus batch size. Solid: total. Dashed: memory time and compute time. The knee is where they cross.
Prefill interrupts decode
The spike every decode user feels
Continuous batching introduces a new failure mode. Prefill is compute-bound and its cost is quadratic in prompt length; decode is memory-bound and its cost is one step. If a newly admitted 8,000-token prompt is prefilled in a single iteration, that iteration cannot also emit the usual token for everyone else — the batch is busy for the whole prompt. Every decoding sequence in the batch sees its inter-token latency jump from one step time to the full prefill time. On a 20 ms decode cadence, an 8,000-token prefill is a 400 ms gap: the user-visible stutter.
The same iteration also delays every other waiting request, so one large prompt creates a queue-wide latency spike. Engines therefore do not schedule prefill and decode in the same rigid way; the next section changes how much prefill goes into one iteration.
Chunked prefill and the token budget
Slice the prompt, pay in TTFT
Chunked prefill caps the number of prefill tokens an iteration may process — vLLM's max_num_batched_tokens. A long prompt is split across several iterations, and the decoding sequences share each of those iterations, so no single step contains a full prompt. The ITL spike becomes a short series of small bumps instead of one cliff. The price is paid by the newcomer: its prefill now takes several iterations, so its time-to-first-token gets worse, not better. The smaller the chunk, the flatter everyone else's ITL and the longer the newcomer waits.
That is the whole trade, and it is explicit below. Toggle chunking on, watch the tallest bar drop and the newcomer's TTFT rise; then shrink the budget and watch both effects intensify. Choosing a budget is choosing a point on that curve — there is no setting that improves both.
Per-step latency seen by an existing decoder. An 8,000-token prefill arrives at step 10; bars show how long that step takes.
Piggybacking
Make the decode batch compute-bound on purpose
Once prefill is chunked it can be piggybacked: a chunk of new prompt tokens rides along inside a decode iteration, and the iteration's tensor work is a mixed batch of prefill rows and decode rows. This is not a hack, it is the point. A pure decode step is memory-bound and leaves the tensor cores idle; a pure prefill step is compute-bound and leaves bandwidth idle. Overlaying them uses both, and under a fixed token budget the scheduler chooses how to split the budget between prompt tokens and decode tokens. Sarathi-Serve's formulation is the cleanest statement of it: prefill chunks and decodes are stitched into one batch per iteration, bounded by a token budget, which keeps the GPU saturated while bounding the worst-case ITL.
Piggybacking is why the token budget is a first-class knob rather than an implementation detail. It also explains a scheduler's preference order: admit as much prefill as fits without starving decode, because the decode tokens were going to be processed anyway and their weight-stream is already being paid for.
Tuning max_num_seqs and max_num_batched_tokens
Two knobs, one frontier
max_num_seqs is the concurrency cap: how many sequences may be resident in a batch. Raise it and you raise throughput until the KV cache runs out or the step turns compute-bound; the V1 scheduler uses a default of 1,024 for small models and 256 for large ones, but the right value is where your p99 stops improving. max_num_batched_tokens is the per-iteration token budget that splits between prefill and decode. Raise it and prompts finish faster, at the cost of larger steps and worse ITL for everyone already decoding.
The two knobs do not act independently: a large token budget with a small sequence cap wastes the budget on a few long prompts, while a small budget with a high cap produces many tiny iterations and high launch overhead. This is exactly the throughput-latency frontier Part 3 draws. The tuner below runs a real iteration-level simulation, sweeps the concurrency cap, and plots each resulting operating point as p99 time-to-first-token against throughput. Move the two sliders and the green point walks the frontier; the faded points are where you have already been.
Lower-left is better. The line is the Pareto frontier over the concurrency sweep; ghost points trace your tuning path.
Cheat sheet
| Policy / knob | What it does | Cost |
|---|---|---|
| Static batching | Fixed group runs to completion together | Padding tax plus a tail where the batch dimension drains |
| Dynamic batching | Dispatch when full or when a queue timer fires | Still lockstep; helps latency at low load only |
| Iteration-level (continuous) | Rebuild the batch every decode step | Scheduler runs every token; prefill can stall the batch |
| Batch size | Amortises the weight read across sequences | Nothing until the memory/compute knee, linear step time after |
| Chunked prefill | Caps prefill tokens per iteration | Flattens ITL spikes; raises the newcomer's TTFT |
| Piggybacking | Runs prefill chunks inside decode iterations | Needs a token budget and a scheduler that can mix the two |
max_num_seqs | Concurrency cap | Throughput up to the KV/step-time limit, then latency |
max_num_batched_tokens | Per-iteration token budget | Prompt latency down, ITL up |
Further reading
- Yu, Jeong, Kim, Kim, Chun, Orca: A Distributed Serving System for Transformer-Based Generative Models, OSDI 2022.
- Agrawal, Kedia, Panwar, Mohan, Kwatra, Gulavani, Tumanov, Ramjee, Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve, OSDI 2024 — chunked prefill and piggybacking.
- Kwon, Li, Zhuang, Sheng, Zheng, Yu, Gonzalez, Zhang, Stoica, PagedAttention, SOSP 2023 — the memory manager continuous batching depends on.
- vLLM documentation, optimization and tuning (
max_num_seqs,max_num_batched_tokens).