Speculative decoding
Decode produces one token per full sweep of the weights. Speculation breaks that rule: a cheap model guesses several tokens, and the expensive model checks all of them in a single pass. Done right it is exactly lossless, and done at the wrong batch size it is a regression.
Verify many, generate one
The inversion
Autoregressive decoding is a sequence of forward passes that each produce a single token. Every pass reads the entire model, so at batch size 1 the step time is essentially weights-over-bandwidth no matter how small the useful output is. Speculative decoding inverts the ratio. A small draft model proposes the next k tokens cheaply; the target model then runs one forward pass over all k proposals at once, in parallel, and accepts the longest prefix that matches what it would itself have generated. When the draft is good, one expensive pass yields several tokens, and the per-token cost of the expensive model falls.
Press play. The draft proposes five tokens in one cheap pass; the target verifies them left to right. The first proposal that disagrees is struck through, the rest are discarded, and the target's own token is substituted at that position — then the cycle begins again from a longer context.
Why the verify pass is nearly free
Memory-bound, not compute-bound
Verifying k tokens is a forward pass in which the query dimension is k instead of 1 — structurally a tiny prefill. The weights are still read exactly once, so the memory time barely changes; only the FLOPs scale with k. For an 8B model on an H100-class part, the decode step is about 4.8 ms of memory time against 0.02 ms of compute for one token. Verifying five tokens raises the compute term to roughly 0.08 ms and leaves the 4.8 ms memory term untouched. The target model's verification is, to first order, free per extra token.
What is not free is the draft. Running a second model, or extra heads, costs real forward passes on the same GPU, and at large batch those passes compete for compute with the requests you are trying to serve. That asymmetry — verification is memory, drafting is compute — is the entire tuning problem of this part, and it is why the batch-size crossover in section 8 exists.
Rejection sampling preserves the distribution
Why this is not approximate
The training guide named speculative decoding but did not open it up (Part 11 of that series lists it among further serving techniques). The important property is that it is exact: the tokens it produces are drawn from the target model's distribution, not an approximation of it. The mechanism is rejection sampling. The draft samples a token x from its distribution q. The target computes its own probability p(x) and accepts with probability min(1, p(x)/q(x)). If the proposal is rejected, a replacement is drawn from the residual distribution proportional to max(0, p(x) − q(x)), renormalized. Some algebra shows the accepted-token distribution is exactly p, so with greedy or sampled target decoding alike the output is identical to what the target alone would have produced — not just similar. The draft's only effect is on speed, never on what the model says.
Acceptance rate, expected tokens, and the optimum k
The formula that sizes everything
Let α be the per-token acceptance probability, assumed roughly constant across positions. The expected number of tokens advanced by one verify pass is a geometric series:
At α = 1 this is k+1; at α = 0 it is 1, and speculation buys nothing. The net speedup is the expected tokens divided by the real cost of producing them, and that denominator grows with k because the draft model must run k cheap steps to make k proposals. With a draft cost c per proposed token, expressed as a fraction of one target step, the net speedup is E[tokens] / (1 + k·c). That denominator is what puts a peak in the curve: past the optimum, extra proposals are more likely to be rejected than to be accepted, and each one still costs a draft step.
Blue (left axis): net speedup. Red (right axis): expected accepted tokens. The dashed line marks the optimum k.
Draft models, Medusa heads, EAGLE
Three ways to get proposals
A separate draft model is the simplest and the original formulation: a smaller model from the same family, served alongside the target. It is exactly lossless and easy to reason about, but it costs a second set of weights in memory and a second forward pass per cycle, which is why it is usually paired with a low k.
Medusa avoids the second model by attaching several lightweight prediction heads to the frozen target backbone, each trained to predict a token at a different offset. The heads can be combined into a tree of candidates verified in one pass, reporting >2.2× for Medusa-1 and 2.3–3.6× for the jointly fine-tuned Medusa-2 — at the cost of a fine-tuning step and extra heads.
EAGLE pushes further by drafting in the target's feature space rather than its token space: a light autoregressive head over hidden states, which raises acceptance because it is conditioned on the target's own internal representation. EAGLE-2 adds a context-aware dynamic draft tree and reports 3.05–4.26× end-to-end, 20–40% over EAGLE-1. EAGLE-3 moves to direct token prediction with multi-layer feature fusion, claims up to 6.5× at low batch and about 1.4× over EAGLE-2, and reports 1.38× throughput even at batch 64 in SGLang. Those maxima are single-request and low-batch numbers; the batch-64 figure is the honest production one.
n-gram and prompt-lookup: a free draft
Zero model, zero parameters
The cheapest draft of all is a lookup. Instead of a model, scan the prompt (and the tokens generated so far) for a suffix that matches the current context, and propose the tokens that followed it last time. There is no draft model, no extra weights, and no draft forward pass — the proposals are nearly free. On workloads where the output copies the input — summarisation, retrieval-augmented answers, code editing, format conversion — the hit rate is high and this is a large, essentially free win. On open-ended chat it is close to useless and can even be a small loss once the verification and bookkeeping overhead are counted, which is why engines expose it as a per-request opt-in rather than a global default. This is the clearest case of speculation being a workload decision before it is a model decision.
Tree attention: verifying a branching draft
One pass, many candidate continuations
A single linear draft chain wastes its budget: if the second token is wrong, the rest of the chain is dead. Instead, draft a tree — at each position, several plausible tokens with their own continuations. The target verifies the entire tree in one forward pass, and the accepted path can branch in a different direction than a linear draft could. The mechanism that makes this valid is a tree attention mask: each candidate token attends to the shared root and to its own ancestors only, and never to tokens in sibling branches. In the flattened token order that mask is block-lower-triangular — attention is confined within each branch's path, which is exactly the block-diagonal structure shown below.
Top: the drafted candidate tree. Bottom: the attention mask — 1 (dark) where a token may attend.
The batch-size crossover
Why engines disable speculation under load
Speculation trades compute for latency. It only pays while decode is memory-bound, which is true at low batch: the verify pass reads the weights once and the extra FLOPs are hidden. As batch grows, the same step becomes compute-bound, and now the draft model's forward passes are stealing tensor-core time from requests that were already saturating the GPU. The two throughput curves cross: speculation wins at low batch and loses above the crossover. A serving engine therefore runs speculation when the queue is short and turns it off (or drops k toward 1) as load rises — the same request gets a latency win when the system is idle and does not tax the system when it is full.
Throughput versus batch size for the target alone and with speculation; they cross once.
Tuning: measure acceptance, not vibes
The production checklist
The one number to instrument is the acceptance rate, broken down by workload and by position in the draft. A high average that comes from the first one or two positions with a near-zero tail is normal; if the tail is collapsing, shorten the draft. Cap k — the formula's optimum is usually between 3 and 7 even at excellent acceptance, because the draft-cost term eventually dominates. Choose the mechanism by workload: n-gram for copy-heavy traffic, a separate draft model when you cannot fine-tune, Medusa/EAGLE when you can and want the last few multiples. And make speculation a load-aware policy, not a deployment constant.
| Method | Acceptance | Reported speedup | Notes |
|---|
Cheat sheet
Recap
| Concept | What it means in practice |
|---|---|
| Draft and verify | Cheap model proposes k tokens; target verifies all in one memory-bound pass and accepts the longest correct prefix |
| Losslessness | Rejection sampling makes the output distribution exactly the target's; speculation is a speed knob only |
| Expected tokens | (1 − αk+1) / (1 − α); at α = 0.8 and k = 3 it is about 2.95 tokens per verify pass |
| Optimum k | Where the marginal acceptance falls below the marginal draft cost; typically 3–7, lower for expensive drafts |
| Tree attention | A branching draft tree verified in one pass with a mask confining each token to its own ancestors |
| Batch crossover | Speculation helps at low batch and hurts once the step is compute-bound — disable it under load |
| Workload fit | n-gram/prompt lookup for copy-heavy output; learned drafts (Medusa, EAGLE) for open-ended generation |
Further reading
References
- Leviathan, Kalman & Matias, Fast Inference from Transformers via Speculative Decoding, ICML 2023 — the exact rejection-sampling argument.
- Chen et al., Accelerating Large Language Model Decoding with Speculative Sampling, 2023.
- Cai et al., Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads, 2024.
- Li et al., EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees, 2024.
- Li et al., EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test, 2025.
- Saxena, Prompt Lookup Decoding — the n-gram draft.