Samplers and schedulers
A diffusion model is trained to denoise one step at a time, but the sampler that walks from noise to image is not fixed by the training. The reverse loop can be stochastic or deterministic, the step size can be constant or warped, and the number of network evaluations is a budget you choose. This part keeps the model fixed and treats the sampler as the object of study: two solver families, the error each one leaves on a known trajectory, and the schedule that decides where the steps land.
The two families
Ancestral noise or a deterministic integration
Every sampler here runs the same reverse loop. At each step it asks the network for an estimate of the noise $\varepsilon$ left in the current sample and then takes a step toward less noise. The families differ in what the step looks like. The ancestral sampler — DDPM — treats the reverse process as a chain of noisy transitions: at every step it computes a mean and then adds a fresh draw of Gaussian noise whose variance is fixed by the schedule. Two runs from the same start give two different images, and the stochasticity is not a defect; it is the Markov chain doing exactly what it was trained to do.
The deterministic family throws the injected noise away. DDIM replaces the noisy transition with a straight interpolation toward the predicted clean sample, so the whole trajectory becomes a function of the starting noise: one start, one endpoint. The same rule can be read as an ordinary differential equation, the probability-flow ODE, whose velocity field is built from the denoiser. Once generation is an ODE, the entire numerical-analysis toolkit applies. Euler is the crudest integrator, Heun corrects Euler with a second field evaluation, and the multistep methods reuse predictions from earlier steps to buy accuracy at a fixed number of network calls.
Solver order and error
What a second evaluation buys
The plot below integrates the probability-flow ODE on a linear schedule from a known start toward a known clean point $x_0 = 1$, and measures how far the endpoint lands from that point. The predictor is deliberately imperfect: it is an oracle model with its noise estimate inflated by $15\%$ whenever the signal has mostly been destroyed, which is the garden-variety failure of a real network at high noise. That single flaw is enough to make the solver's job non-trivial, and it is what the curves measure.
Euler is first order: halve the step size and the error roughly halves. Heun is second order: it takes an Euler step, evaluates the field again at the provisional endpoint, and averages the two, which costs twice as many network calls per step but makes the error fall much faster as the steps multiply. The dashed line is the floor the error would reach with a perfect predictor, computed by Diffusion.solverError; no solver can go below it.
Endpoint error against the number of steps, log scale. Solid curves use the imperfect predictor; the dashed curve is the perfect-model floor.
Timestep spacing
Where the steps land on the sigma curve
The schedule fixes how much noise each timestep carries. It is convenient to measure that as a noise-to-signal ratio, $\sigma_t = \sqrt{(1-\bar\alpha_t)/\bar\alpha_t}$, which falls from a large value at the start of sampling to nearly zero at the end. Uniform spacing — the familiar linspace over timesteps — takes equal steps along the index, but the index is not where the difficulty lives. Almost all of the structural decision-making happens when $\sigma$ is small and the image is nearly formed, and uniform spacing spends most of its budget at high $\sigma$ where there is little to decide.
Karras and coauthors warp the grid in $\sigma$-space instead, pulling steps densely toward small $\sigma$ and sparsely toward the noise end. The plot makes the difference visible: the marker sets are the same number of steps, but the Karras set leaves the smooth $\sigma$ curve and clusters at the low-noise end. With a fixed step budget, that reallocation is usually worth more than a higher-order solver, which is why every production sampler exposes a spacing knob beside its solver knob.
The noise-to-signal curve (solid) with the steps chosen by uniform spacing and by Karras spacing. Both use the step count from the solver demo above.
Larger ρ concentrates steps even harder at low σ. Uniform spacing is the special case where every step covers the same index interval.
The step budget
The twenty-to-fifty sweet spot
The training process used a thousand forward steps, but sampling never has to. The reverse loop is far more forgiving than the forward chain, and the honest way to state the trade is in network evaluations: wall-clock time is roughly the number of forward passes times the cost of one pass. The bars below hold the solver fixed and spend four different budgets, from four steps to fifty, on the same trajectory.
In practice the curve flattens somewhere in the twenties. Below that the image is visibly under-resolved and the sampler has not had enough looks at the low-noise end; above fifty the extra evaluations buy a difference no one can see. That flat region is why the default sampler in a shipping pipeline is typically twenty to fifty steps with a second-order or multistep solver and a warped grid, and why the few-step methods in the next part are interesting at all: they try to move the knee of this curve down to one or four.
Endpoint error at 4, 8, 20 and 50 steps for the solver selected above. The highlighted bar is the current budget.
Multistep solvers such as DPM-Solver++ reach the flat region in roughly ten to twenty calls by reusing earlier predictions as extra terms in the update, rather than by evaluating the field again.
Where this shows up
The knobs every pipeline exposes
Stable Diffusion samplers
The sampler menu in every Stable Diffusion interface is this part's taxonomy: ancestral samplers that inject noise, DDIM and its relatives that do not, and DPM-Solver++ variants that add history terms. The reverse-loop chapter builds the same loop on a real trained denoiser.
Few-step and interactive generation
A straight-path flow-matching field can be integrated in a handful of Euler steps, which is what makes interactive editing and robot policies at control rate feasible. The ODE guide develops Euler and Runge–Kutta error orders, which is the analysis behind every curve on this page.
The same arithmetic organises video, audio and 3-D generation, where each network call is even more expensive and the step budget is measured in seconds of latency rather than milliseconds. Whenever a paper reports a sampling time, read it as solver, spacing and step count together — the three are chosen jointly and usually tuned against each other.
Further reading
Sampling is where diffusion stops being a probabilistic model and starts being numerical analysis. These are the papers that made that shift explicit: the one that removed the injected noise, the one that added higher-order integration, and the one that treated the schedule as a warped grid in noise-to-signal space.
Song and coauthors show that the reverse process has a deterministic counterpart; Lu and coauthors bring the multistep ODE solvers of numerical analysis into the loop; Karras and coauthors rework the whole sampler around the sigma parameterisation; and Ho and coauthors are the ancestral baseline everything else is measured against.
- Jiaming Song, Chenlin Meng and Stefano Ermon, "Denoising Diffusion Implicit Models", 2021 — the deterministic sampler and the step-count trade.
- Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li and Jun Zhu, "DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps", 2022 — high-order and multistep solvers for the reverse ODE.
- Tero Karras, Miika Aittala, Timo Aila and Samuli Laine, "Elucidating the Design Space of Diffusion-Based Generative Models", 2022 — the sigma parameterisation and the non-uniform grid.
- Jonathan Ho, Ajay Jain and Pieter Abbeel, "Denoising Diffusion Probabilistic Models", 2020 — the ancestral sampler in its original form.
Cheat sheet
| Term | Meaning here |
|---|---|
| Ancestral sampler | Injects fresh Gaussian noise at every step; DDPM; stochastic output |
| Deterministic sampler | No injected noise; the endpoint is a function of the starting noise |
| Probability-flow ODE | The differential equation whose solution is the deterministic reverse path |
| Euler step | First order: follow the field at the current point; $x \leftarrow x + v\,\Delta t$ |
| Heun step | Second order: average the field at the start and at a provisional endpoint; two evaluations |
| NFE | Number of function evaluations; the true cost of a sampler |
| $\sigma_t$ | Noise-to-signal ratio $\sqrt{(1-\bar\alpha_t)/\bar\alpha_t}$; what the schedule really controls |
| Karras spacing | Non-uniform grid, dense at low $\sigma$; Diffusion.karrasSigmas |
| Sweet spot | Roughly 20–50 evaluations for images before the error curve flattens |