Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

One-step generation

Replace the trajectory with a single map

A deterministic sampler traces a trajectory from a noise sample to a clean sample, and the trajectory is a function: name a starting noise $\varepsilon$ and the sampler returns an image. Collapsing fifty steps into one means learning that whole function at once. The student takes the noisy input at some time $t$ and predicts the clean sample directly, so a single forward pass substitutes for the rest of the loop.

The target is available for free. Run the slow teacher forward from a clean example to get its noisy version $x_t$, and the label is the clean example $x_0$ the teacher started from; the student regresses $x_t \mapsto x_0$ over many times $t$. This is the consistency objective in its simplest form, and it is exactly what the demo below trains. The teacher is an oracle in this page so that the arithmetic stays honest, but the structure is the same one that turns Stable Diffusion into a four-step model.

💡 By the end of this part you'll be able to name the four families of few-step methods, say what each one distils into the student, and explain why every one of them trades sample diversity for the step count it removes.
2

Training the student

A real network, trained here, in one page load

The target distribution is the three-cluster mixture used elsewhere in this series: a small set of 2-D points drawn from Diffusion.toy2d. The student is a tiny multilayer perceptron with two inputs and two outputs, trained on pairs built by Diffusion.distillPair — take a clean point, pick a time, draw the noisy version, and ask the network for the clean point back. The scatter below shows the result of sampling that student with the current step budget.

Move the step slider and watch two things at once. The student's samples start as a tight blob near the origin, because when the input carries almost no signal the best prediction is the average of the whole mixture. As the step count grows, the sampling loop feeds the student progressively cleaner inputs, its predictions commit to a cluster, and the samples spread out over the three modes. A one-step student has to make that commitment from pure noise, and it cannot.

Grey: the target mixture, and the endpoints of the fifty-step oracle teacher trajectory. Pink: samples from the trained student at the selected step count.

A one-step generator is a single forward pass. Four steps is the practical target most production few-step models aim for.

⚠ Mode averaging is the failure the picture is named after. Asked for the clean sample behind a nearly noise-free input, the squared-error-optimal answer is the conditional mean of the data, which for a multimodal distribution is a location where no data actually live. The student is being honest; the loss is what makes it blurry.
3

The distillation families

Four ways to move the knowledge

Progressive distillation halves the count repeatedly. Train a two-step model to match one step of a four-step teacher, then a one-step model to match one step of the two-step model, each stage supervised by the teacher's own trajectories rather than by the data. The division of labour is the point: every stage is an easier approximation than the last, and the student never has to learn a fifty-to-one jump.

Consistency models and their latent version, the Latent Consistency Model, enforce a single condition instead of a staged one: the map from any point on a trajectory to the trajectory's endpoint must be the same, whether the point is noisy or clean. This self-consistency makes the student a one-step generator by construction, and few-step sampling is just applying the map with a couple of re-noise rounds. The student in this page is the toy version of that idea.

Adversarial distillation — ADD and the distribution-matching variant DMD — drops the regression target and keeps a discriminator or a score-matching term instead. Instead of asking the student to reproduce any single teacher trajectory, it asks the student's whole output distribution to look like the teacher's, which recovers diversity that a squared-error objective averages away. Shortcut and flow-map methods fold the trajectory itself: a shortcut model predicts not the endpoint but a step of a chosen size, and a flow map composes those steps, so the number of evaluations becomes a dial the student was trained to honour.

💡 The families differ in what they supervise: a trajectory (progressive), a fixed point (consistency), a distribution (adversarial), or a step-size-parameterised map (shortcut). Choosing between them is choosing which property of the teacher you are unwilling to lose.
4

What the speed costs

Fifty steps into one

The top plot below is the diversity readout made visible: the spread of the student's samples against the step count, with the teacher's spread as a horizontal reference. Every step removed narrows the cloud. Collapsing the whole trajectory into one evaluation is a big win in latency and a real loss in variety, and the loss is not an artefact of this toy — it is the same averaging that makes a one-step image generator produce a slightly plastic, over-smoothed look.

The bottom canvas is the trajectory view of the same idea. A fifty-step teacher walks a curved path down to the clean point; a four-step student approximates that path with four chords; a one-step student replaces the curve with a single straight jump. Each approximation is faithful at the endpoints and coarse in between, which is why the shortest few-step schedules are the ones most sensitive to the schedule's shape.

Top: sample spread against step count, with the fifty-step teacher as the dashed reference. Bottom: the same trajectory integrated with 50, 4 and 1 steps.

Diversity is measured as the standard deviation of the generated cloud; coverage is the fraction of target modes that receive at least one sample.

5

Where this shows up

The models built for a latency budget

Image

LCM, Turbo and Lightning

Latent Consistency Models, SDXL-Turbo and SDXL-Lightning are three answers to the same latency budget, built respectively on the consistency condition, adversarial distillation and progressive distillation. All three are evaluated with the sampler machinery from the previous part, and all three expose a trade between steps and diversity.

Flow & video

Schnell and real-time video

Flux's schnell variant uses a rectified-flow student distilled for four steps, and the same recipe is what makes interactive video and control-rate robot policies feasible. Adversarial and flow-map variants appear where the diversity loss is most visible — in video, where a collapsed sample reads as a frozen scene.

The pattern recurs whenever a generative model is deployed: a slow, faithful teacher for quality and a distilled student for throughput, served side by side. The guidance chapter shows the other half of the same trade, where a scale knob moves quality against diversity at a fixed step count.

Further reading

These papers are best read as four attempts at the same problem: how to make one network evaluation stand in for many. Start with the staged method, then the fixed-point method, then the distributional and map-based corrections that recover diversity.

Salimans and Ho staged the collapse; Song and coauthors replaced the staging with a self-consistency condition; Luo and coauthors moved that condition into the latent space; Sauer and coauthors added a discriminator; and Frans and coauthors generalised the whole thing to a step-size dial.

Cheat sheet

TermMeaning here
TeacherThe slow, faithful sampler whose trajectory the student imitates
One-step studentPredicts $x_0$ from $x_t$ in a single forward pass, for any $t$
Progressive distillationHalve the step count repeatedly; each stage imitates the previous
Consistency modelEnforce one map from every point of a trajectory to its endpoint
LCMConsistency enforced in the latent space of a diffusion autoencoder
ADD / DMDAdversarial or distribution-matching objectives that recover diversity
Shortcut / flow mapPredict a step of a chosen size; compose steps to spend a chosen budget
Diversity costFewer steps narrows the output cloud; regression averages modes away
7

Check your understanding

0/4 answered