Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Static batching and the padding tax

One batch, one finish line

The oldest way to batch, and the one almost every other accelerator workload uses, is static batching: collect N requests, stack them, run the batch until every member has finished, then start the next batch. Matrix kernels need rectangular inputs, so the short sequences are padded out to the length of the longest one and the extra positions masked. Masking the arithmetic does not remove it: padded tokens cost the same FLOPs and the same attention reads as real ones. This is the padding tax, and it is largest exactly where chatbots live, in requests whose output lengths vary by an order of magnitude.

The second cost is worse than the padding. A batch cannot release a slot until its slowest member finishes, so a single 2,000-token response holds five finished 100-token responses hostage. The GPU's batch dimension — the thing that amortises a memory-bound decode step across requests — sits partly empty for the whole tail of every batch. Drag the variance slider below to see the tax itself; the Gantt in section 3 shows what it does to the timeline.

Six sequences padded to the longest in the batch. Hatched cells are masked padding that still cost compute.

2

Dynamic batching

Stop waiting for a full batch

Dynamic batching keeps the lockstep execution but changes when a batch is dispatched. Instead of waiting for exactly N requests, the scheduler starts a batch as soon as either N have queued or a maximum queue delay has elapsed, whichever comes first. At low load this dispatches small batches immediately, which is a large latency win; at high load it behaves like static batching, because N arrives long before the timer fires.

What it does not change is the finish line. The batch still runs to completion together, so padding and tail waste remain, and the GPU still idles between batches. Dynamic batching is the right tool for vision models with near-constant latency; for autoregressive decoding, where request durations differ by 100×, it is a half-measure. The demo below lets the same workload run under both regimes and under iteration-level batching.

3

Iteration-level batching (Orca)

The batch is rebuilt every token

Orca's observation, in 2022, was that a transformer decode step is already an iteration boundary: the model produces one token per sequence, and then the scheduler is free to change who is in the batch. Iteration-level batching (equivalently continuous or in-flight batching) does exactly that. Finished sequences leave the batch the moment they emit their stop token; waiting sequences join the moment the KV pool and the sequence cap allow. The batch dimension stays nearly full for the entire run, and no request waits for a stranger's last token.

The training guide's continuous batching demo shows the mechanism on a handful of requests. What it does not show is what the scheduler is optimising — that is this part's job — nor the two knobs that keep the iteration from becoming either idle or overstuffed, nor the prefill problem that arrives with it. The Gantt chart below runs one fixed request stream through all three policies.

Static batching
Dynamic batching (2.5 s dispatch window)
Continuous / iteration-level batching
💡 Read the three axes. Grey is queued, blue is prefill, the darker band is decode. Utilisation is the fraction of the run in which at least one sequence was computing; throughput is finished output tokens per second over the whole run; p99 e2e is the worst-request latency.
4

What batch size does to step time

Flat, then linear, then you pay

Part 1's memory wall says a decode step reads the weights once per token no matter how large the batch. So adding the second sequence to a decode batch is almost free: the same 70 GB of weights streams from HBM while the tensor cores process two rows instead of one. This is why decode throughput scales nearly linearly with batch at small batch sizes. It stops when the batch's arithmetic intensity reaches the machine's ridge point — the batch size at which 2 × params × batch FLOPs take longer than streaming weights + KV bytes. Past that knee the step is compute-bound and its time grows linearly in batch.

The knee is a property of the machine and the model, not a tuning decision: for an 8B model at BF16 on H100 it lands near batch 295 with no cache, and the KV cache drags it down toward the batch sizes people actually use. The step time itself — the quantity that becomes the inter-token latency — is what the chart plots. Slide batch, and slide the cache per sequence, and watch the knee move.

Decode step time versus batch size. Solid: total. Dashed: memory time and compute time. The knee is where they cross.

5

Prefill interrupts decode

The spike every decode user feels

Continuous batching introduces a new failure mode. Prefill is compute-bound and its cost is quadratic in prompt length; decode is memory-bound and its cost is one step. If a newly admitted 8,000-token prompt is prefilled in a single iteration, that iteration cannot also emit the usual token for everyone else — the batch is busy for the whole prompt. Every decoding sequence in the batch sees its inter-token latency jump from one step time to the full prefill time. On a 20 ms decode cadence, an 8,000-token prefill is a 400 ms gap: the user-visible stutter.

The same iteration also delays every other waiting request, so one large prompt creates a queue-wide latency spike. Engines therefore do not schedule prefill and decode in the same rigid way; the next section changes how much prefill goes into one iteration.

6

Chunked prefill and the token budget

Slice the prompt, pay in TTFT

Chunked prefill caps the number of prefill tokens an iteration may process — vLLM's max_num_batched_tokens. A long prompt is split across several iterations, and the decoding sequences share each of those iterations, so no single step contains a full prompt. The ITL spike becomes a short series of small bumps instead of one cliff. The price is paid by the newcomer: its prefill now takes several iterations, so its time-to-first-token gets worse, not better. The smaller the chunk, the flatter everyone else's ITL and the longer the newcomer waits.

That is the whole trade, and it is explicit below. Toggle chunking on, watch the tallest bar drop and the newcomer's TTFT rise; then shrink the budget and watch both effects intensify. Choosing a budget is choosing a point on that curve — there is no setting that improves both.

Per-step latency seen by an existing decoder. An 8,000-token prefill arrives at step 10; bars show how long that step takes.

7

Piggybacking

Make the decode batch compute-bound on purpose

Once prefill is chunked it can be piggybacked: a chunk of new prompt tokens rides along inside a decode iteration, and the iteration's tensor work is a mixed batch of prefill rows and decode rows. This is not a hack, it is the point. A pure decode step is memory-bound and leaves the tensor cores idle; a pure prefill step is compute-bound and leaves bandwidth idle. Overlaying them uses both, and under a fixed token budget the scheduler chooses how to split the budget between prompt tokens and decode tokens. Sarathi-Serve's formulation is the cleanest statement of it: prefill chunks and decodes are stitched into one batch per iteration, bounded by a token budget, which keeps the GPU saturated while bounding the worst-case ITL.

Piggybacking is why the token budget is a first-class knob rather than an implementation detail. It also explains a scheduler's preference order: admit as much prefill as fits without starving decode, because the decode tokens were going to be processed anyway and their weight-stream is already being paid for.

8

Tuning max_num_seqs and max_num_batched_tokens

Two knobs, one frontier

max_num_seqs is the concurrency cap: how many sequences may be resident in a batch. Raise it and you raise throughput until the KV cache runs out or the step turns compute-bound; the V1 scheduler uses a default of 1,024 for small models and 256 for large ones, but the right value is where your p99 stops improving. max_num_batched_tokens is the per-iteration token budget that splits between prefill and decode. Raise it and prompts finish faster, at the cost of larger steps and worse ITL for everyone already decoding.

The two knobs do not act independently: a large token budget with a small sequence cap wastes the budget on a few long prompts, while a small budget with a high cap produces many tiny iterations and high launch overhead. This is exactly the throughput-latency frontier Part 3 draws. The tuner below runs a real iteration-level simulation, sweeps the concurrency cap, and plots each resulting operating point as p99 time-to-first-token against throughput. Move the two sliders and the green point walks the frontier; the faded points are where you have already been.

Lower-left is better. The line is the Pareto frontier over the concurrency sweep; ghost points trace your tuning path.

💡 The tuner's punchline. Both knobs move the same operating point. Before tuning either, decide which axis your SLO is written against — TTFT or ITL — because every step along this frontier makes one of them worse.

Cheat sheet

Policy / knobWhat it doesCost
Static batchingFixed group runs to completion togetherPadding tax plus a tail where the batch dimension drains
Dynamic batchingDispatch when full or when a queue timer firesStill lockstep; helps latency at low load only
Iteration-level (continuous)Rebuild the batch every decode stepScheduler runs every token; prefill can stall the batch
Batch sizeAmortises the weight read across sequencesNothing until the memory/compute knee, linear step time after
Chunked prefillCaps prefill tokens per iterationFlattens ITL spikes; raises the newcomer's TTFT
PiggybackingRuns prefill chunks inside decode iterationsNeeds a token budget and a scheduler that can mix the two
max_num_seqsConcurrency capThroughput up to the KV/step-time limit, then latency
max_num_batched_tokensPer-iteration token budgetPrompt latency down, ITL up

Further reading

?

Check your understanding

0/5 answered