Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Verify many, generate one

The inversion

Autoregressive decoding is a sequence of forward passes that each produce a single token. Every pass reads the entire model, so at batch size 1 the step time is essentially weights-over-bandwidth no matter how small the useful output is. Speculative decoding inverts the ratio. A small draft model proposes the next k tokens cheaply; the target model then runs one forward pass over all k proposals at once, in parallel, and accepts the longest prefix that matches what it would itself have generated. When the draft is good, one expensive pass yields several tokens, and the per-token cost of the expensive model falls.

Press play. The draft proposes five tokens in one cheap pass; the target verifies them left to right. The first proposal that disagrees is struck through, the rest are discarded, and the target's own token is substituted at that position — then the cycle begins again from a longer context.

2

Why the verify pass is nearly free

Memory-bound, not compute-bound

Verifying k tokens is a forward pass in which the query dimension is k instead of 1 — structurally a tiny prefill. The weights are still read exactly once, so the memory time barely changes; only the FLOPs scale with k. For an 8B model on an H100-class part, the decode step is about 4.8 ms of memory time against 0.02 ms of compute for one token. Verifying five tokens raises the compute term to roughly 0.08 ms and leaves the 4.8 ms memory term untouched. The target model's verification is, to first order, free per extra token.

What is not free is the draft. Running a second model, or extra heads, costs real forward passes on the same GPU, and at large batch those passes compete for compute with the requests you are trying to serve. That asymmetry — verification is memory, drafting is compute — is the entire tuning problem of this part, and it is why the batch-size crossover in section 8 exists.

3

Rejection sampling preserves the distribution

Why this is not approximate

The training guide named speculative decoding but did not open it up (Part 11 of that series lists it among further serving techniques). The important property is that it is exact: the tokens it produces are drawn from the target model's distribution, not an approximation of it. The mechanism is rejection sampling. The draft samples a token x from its distribution q. The target computes its own probability p(x) and accepts with probability min(1, p(x)/q(x)). If the proposal is rejected, a replacement is drawn from the residual distribution proportional to max(0, p(x) − q(x)), renormalized. Some algebra shows the accepted-token distribution is exactly p, so with greedy or sampled target decoding alike the output is identical to what the target alone would have produced — not just similar. The draft's only effect is on speed, never on what the model says.

💡 Consequence for tuning: because the distribution is preserved, you can change the draft model, tree shape, or k at runtime, per request and per load condition, without changing your quality contract. Every degree of freedom here is purely a performance knob.
4

Acceptance rate, expected tokens, and the optimum k

The formula that sizes everything

Let α be the per-token acceptance probability, assumed roughly constant across positions. The expected number of tokens advanced by one verify pass is a geometric series:

$$E[\text{tokens}] = \frac{1-\alpha^{k+1}}{1-\alpha}$$

At α = 1 this is k+1; at α = 0 it is 1, and speculation buys nothing. The net speedup is the expected tokens divided by the real cost of producing them, and that denominator grows with k because the draft model must run k cheap steps to make k proposals. With a draft cost c per proposed token, expressed as a fraction of one target step, the net speedup is E[tokens] / (1 + k·c). That denominator is what puts a peak in the curve: past the optimum, extra proposals are more likely to be rejected than to be accepted, and each one still costs a draft step.

Blue (left axis): net speedup. Red (right axis): expected accepted tokens. The dashed line marks the optimum k.

5

Draft models, Medusa heads, EAGLE

Three ways to get proposals

A separate draft model is the simplest and the original formulation: a smaller model from the same family, served alongside the target. It is exactly lossless and easy to reason about, but it costs a second set of weights in memory and a second forward pass per cycle, which is why it is usually paired with a low k.

Medusa avoids the second model by attaching several lightweight prediction heads to the frozen target backbone, each trained to predict a token at a different offset. The heads can be combined into a tree of candidates verified in one pass, reporting >2.2× for Medusa-1 and 2.3–3.6× for the jointly fine-tuned Medusa-2 — at the cost of a fine-tuning step and extra heads.

EAGLE pushes further by drafting in the target's feature space rather than its token space: a light autoregressive head over hidden states, which raises acceptance because it is conditioned on the target's own internal representation. EAGLE-2 adds a context-aware dynamic draft tree and reports 3.05–4.26× end-to-end, 20–40% over EAGLE-1. EAGLE-3 moves to direct token prediction with multi-layer feature fusion, claims up to 6.5× at low batch and about 1.4× over EAGLE-2, and reports 1.38× throughput even at batch 64 in SGLang. Those maxima are single-request and low-batch numbers; the batch-64 figure is the honest production one.

6

n-gram and prompt-lookup: a free draft

Zero model, zero parameters

The cheapest draft of all is a lookup. Instead of a model, scan the prompt (and the tokens generated so far) for a suffix that matches the current context, and propose the tokens that followed it last time. There is no draft model, no extra weights, and no draft forward pass — the proposals are nearly free. On workloads where the output copies the input — summarisation, retrieval-augmented answers, code editing, format conversion — the hit rate is high and this is a large, essentially free win. On open-ended chat it is close to useless and can even be a small loss once the verification and bookkeeping overhead are counted, which is why engines expose it as a per-request opt-in rather than a global default. This is the clearest case of speculation being a workload decision before it is a model decision.

7

Tree attention: verifying a branching draft

One pass, many candidate continuations

A single linear draft chain wastes its budget: if the second token is wrong, the rest of the chain is dead. Instead, draft a tree — at each position, several plausible tokens with their own continuations. The target verifies the entire tree in one forward pass, and the accepted path can branch in a different direction than a linear draft could. The mechanism that makes this valid is a tree attention mask: each candidate token attends to the shared root and to its own ancestors only, and never to tokens in sibling branches. In the flattened token order that mask is block-lower-triangular — attention is confined within each branch's path, which is exactly the block-diagonal structure shown below.

Top: the drafted candidate tree. Bottom: the attention mask — 1 (dark) where a token may attend.

8

The batch-size crossover

Why engines disable speculation under load

Speculation trades compute for latency. It only pays while decode is memory-bound, which is true at low batch: the verify pass reads the weights once and the extra FLOPs are hidden. As batch grows, the same step becomes compute-bound, and now the draft model's forward passes are stealing tensor-core time from requests that were already saturating the GPU. The two throughput curves cross: speculation wins at low batch and loses above the crossover. A serving engine therefore runs speculation when the queue is short and turns it off (or drops k toward 1) as load rises — the same request gets a latency win when the system is idle and does not tax the system when it is full.

Throughput versus batch size for the target alone and with speculation; they cross once.

9

Tuning: measure acceptance, not vibes

The production checklist

The one number to instrument is the acceptance rate, broken down by workload and by position in the draft. A high average that comes from the first one or two positions with a near-zero tail is normal; if the tail is collapsing, shorten the draft. Cap k — the formula's optimum is usually between 3 and 7 even at excellent acceptance, because the draft-cost term eventually dominates. Choose the mechanism by workload: n-gram for copy-heavy traffic, a separate draft model when you cannot fine-tune, Medusa/EAGLE when you can and want the last few multiples. And make speculation a load-aware policy, not a deployment constant.

⚠️ Paper maxima are not production numbers. The speedups below are best-case, usually single-request or low-batch, and the n-gram figure in particular is highly workload-dependent and can fall below 1× when the hit rate is low. Treat the ranking as a starting hypothesis and confirm it against your own traces.
MethodAcceptanceReported speedupNotes

✓

Cheat sheet

Recap

ConceptWhat it means in practice
Draft and verifyCheap model proposes k tokens; target verifies all in one memory-bound pass and accepts the longest correct prefix
LosslessnessRejection sampling makes the output distribution exactly the target's; speculation is a speed knob only
Expected tokens(1 − αk+1) / (1 − α); at α = 0.8 and k = 3 it is about 2.95 tokens per verify pass
Optimum kWhere the marginal acceptance falls below the marginal draft cost; typically 3–7, lower for expensive drafts
Tree attentionA branching draft tree verified in one pass with a mask confining each token to its own ancestors
Batch crossoverSpeculation helps at low batch and hurts once the step is compute-bound — disable it under load
Workload fitn-gram/prompt lookup for copy-heavy output; learned drafts (Medusa, EAGLE) for open-ended generation
📚

Further reading

References

?

Check your understanding

0/5 answered