Few-step generation
Sampling cost is the number of times the network is evaluated, and a large denoiser evaluated fifty times is the reason a diffusion image takes seconds. Distillation attacks the count directly: train a second model whose whole job is to reproduce, in one or four evaluations, what the slow model does in fifty. This part trains a real one-step student in your browser, watches its samples collapse toward the mean as the step count falls, and organises the four families of methods that buy the speed.
One-step generation
Replace the trajectory with a single map
A deterministic sampler traces a trajectory from a noise sample to a clean sample, and the trajectory is a function: name a starting noise $\varepsilon$ and the sampler returns an image. Collapsing fifty steps into one means learning that whole function at once. The student takes the noisy input at some time $t$ and predicts the clean sample directly, so a single forward pass substitutes for the rest of the loop.
The target is available for free. Run the slow teacher forward from a clean example to get its noisy version $x_t$, and the label is the clean example $x_0$ the teacher started from; the student regresses $x_t \mapsto x_0$ over many times $t$. This is the consistency objective in its simplest form, and it is exactly what the demo below trains. The teacher is an oracle in this page so that the arithmetic stays honest, but the structure is the same one that turns Stable Diffusion into a four-step model.
Training the student
A real network, trained here, in one page load
The target distribution is the three-cluster mixture used elsewhere in this series: a small set of 2-D points drawn from Diffusion.toy2d. The student is a tiny multilayer perceptron with two inputs and two outputs, trained on pairs built by Diffusion.distillPair — take a clean point, pick a time, draw the noisy version, and ask the network for the clean point back. The scatter below shows the result of sampling that student with the current step budget.
Move the step slider and watch two things at once. The student's samples start as a tight blob near the origin, because when the input carries almost no signal the best prediction is the average of the whole mixture. As the step count grows, the sampling loop feeds the student progressively cleaner inputs, its predictions commit to a cluster, and the samples spread out over the three modes. A one-step student has to make that commitment from pure noise, and it cannot.
Grey: the target mixture, and the endpoints of the fifty-step oracle teacher trajectory. Pink: samples from the trained student at the selected step count.
A one-step generator is a single forward pass. Four steps is the practical target most production few-step models aim for.
The distillation families
Four ways to move the knowledge
Progressive distillation halves the count repeatedly. Train a two-step model to match one step of a four-step teacher, then a one-step model to match one step of the two-step model, each stage supervised by the teacher's own trajectories rather than by the data. The division of labour is the point: every stage is an easier approximation than the last, and the student never has to learn a fifty-to-one jump.
Consistency models and their latent version, the Latent Consistency Model, enforce a single condition instead of a staged one: the map from any point on a trajectory to the trajectory's endpoint must be the same, whether the point is noisy or clean. This self-consistency makes the student a one-step generator by construction, and few-step sampling is just applying the map with a couple of re-noise rounds. The student in this page is the toy version of that idea.
Adversarial distillation — ADD and the distribution-matching variant DMD — drops the regression target and keeps a discriminator or a score-matching term instead. Instead of asking the student to reproduce any single teacher trajectory, it asks the student's whole output distribution to look like the teacher's, which recovers diversity that a squared-error objective averages away. Shortcut and flow-map methods fold the trajectory itself: a shortcut model predicts not the endpoint but a step of a chosen size, and a flow map composes those steps, so the number of evaluations becomes a dial the student was trained to honour.
What the speed costs
Fifty steps into one
The top plot below is the diversity readout made visible: the spread of the student's samples against the step count, with the teacher's spread as a horizontal reference. Every step removed narrows the cloud. Collapsing the whole trajectory into one evaluation is a big win in latency and a real loss in variety, and the loss is not an artefact of this toy — it is the same averaging that makes a one-step image generator produce a slightly plastic, over-smoothed look.
The bottom canvas is the trajectory view of the same idea. A fifty-step teacher walks a curved path down to the clean point; a four-step student approximates that path with four chords; a one-step student replaces the curve with a single straight jump. Each approximation is faithful at the endpoints and coarse in between, which is why the shortest few-step schedules are the ones most sensitive to the schedule's shape.
Top: sample spread against step count, with the fifty-step teacher as the dashed reference. Bottom: the same trajectory integrated with 50, 4 and 1 steps.
Diversity is measured as the standard deviation of the generated cloud; coverage is the fraction of target modes that receive at least one sample.
Where this shows up
The models built for a latency budget
LCM, Turbo and Lightning
Latent Consistency Models, SDXL-Turbo and SDXL-Lightning are three answers to the same latency budget, built respectively on the consistency condition, adversarial distillation and progressive distillation. All three are evaluated with the sampler machinery from the previous part, and all three expose a trade between steps and diversity.
Schnell and real-time video
Flux's schnell variant uses a rectified-flow student distilled for four steps, and the same recipe is what makes interactive video and control-rate robot policies feasible. Adversarial and flow-map variants appear where the diversity loss is most visible — in video, where a collapsed sample reads as a frozen scene.
The pattern recurs whenever a generative model is deployed: a slow, faithful teacher for quality and a distilled student for throughput, served side by side. The guidance chapter shows the other half of the same trade, where a scale knob moves quality against diversity at a fixed step count.
Further reading
These papers are best read as four attempts at the same problem: how to make one network evaluation stand in for many. Start with the staged method, then the fixed-point method, then the distributional and map-based corrections that recover diversity.
Salimans and Ho staged the collapse; Song and coauthors replaced the staging with a self-consistency condition; Luo and coauthors moved that condition into the latent space; Sauer and coauthors added a discriminator; and Frans and coauthors generalised the whole thing to a step-size dial.
- Tim Salimans and Jonathan Ho, "Progressive Distillation for Fast Sampling of Diffusion Models", 2022 — halving the step count stage by stage.
- Yang Song, Prafulla Dhariwal, Mark Chen and Ilya Sutskever, "Consistency Models", 2023 — the fixed-point condition that makes one-step generation possible.
- Simian Luo, Yiqin Tan, Longbo Huang, Jian Li and Hang Zhao, "Latent Consistency Models", 2023 — consistency in a compressed latent space.
- Axel Sauer, Dominik Lorenz, Andreas Blattmann and Robin Rombach, "Adversarial Diffusion Distillation", 2023 — a discriminator that restores the diversity regression removes.
- Kevin Frans, Danijar Hafner, Sergey Levine and Pieter Abbeel, "One Step Diffusion via Shortcut Models", 2024 — step size as a training-time dial.
Cheat sheet
| Term | Meaning here |
|---|---|
| Teacher | The slow, faithful sampler whose trajectory the student imitates |
| One-step student | Predicts $x_0$ from $x_t$ in a single forward pass, for any $t$ |
| Progressive distillation | Halve the step count repeatedly; each stage imitates the previous |
| Consistency model | Enforce one map from every point of a trajectory to its endpoint |
| LCM | Consistency enforced in the latent space of a diffusion autoencoder |
| ADD / DMD | Adversarial or distribution-matching objectives that recover diversity |
| Shortcut / flow map | Predict a step of a chosen size; compose steps to spend a chosen budget |
| Diversity cost | Fewer steps narrows the output cloud; regression averages modes away |