Learning to undo one step
The forward process is a recipe you wrote yourself, so undoing it ought to be easy. It is not, because the information the noise destroyed is gone — all a model can do is estimate what was probably there. This part trains a real neural network in your browser on a two-dimensional toy distribution, watches its loss fall, and uses it to move noisy points back toward the data. Along the way it answers a question that sounds like bookkeeping but is not: should the network predict the noise, the clean point, or something in between?
The denoising objective
Ask the network to name the noise
Take a clean point $x_0$, pick a random timestep $t$, and use the closed form from the previous part to jump straight to $x_t = \sqrt{\bar\alpha_t}\,x_0 + \sqrt{1-\bar\alpha_t}\,\varepsilon$. You know $\varepsilon$ exactly, because you generated it. So the training objective can be as blunt as this: given $x_t$ and $t$, predict $\varepsilon$, and penalise the squared error.
$$ \mathcal{L}(\theta) = \mathbb{E}_{x_0,\,t,\,\varepsilon}\!\left[\, \lVert \varepsilon - \varepsilon_\theta(x_t, t) \rVert^2 \,\right] $$
This is the simplified DDPM loss, and its elegance is that a hard probabilistic problem has become ordinary regression. There is no adversarial game and no likelihood to estimate directly: draw a batch, jump to a random noise level, and take a gradient step. The reason the shortcut is legitimate is that predicting the noise that was added is equivalent to estimating the score of the noisy distribution up to a known factor, a fact the fifth part makes precise.
The network below does not know about images. It is a tiny multilayer perceptron with two inputs, two tanh hidden layers of sixteen units, and two outputs, trained with hand-written backpropagation on a toy distribution — three Gaussian blobs in the plane. Everything true of this network is true of a billion-parameter U-Net; the only difference is scale.
Train a denoiser in the browser
A real network, a real loss curve, one slider
The grey dots are the toy dataset: three clusters, roughly two hundred points. The blue dots are those points after the forward process at the noise level the slider selects, so at low $t$ they sit on the data and at high $t$ they are scattered almost at random. The pink arrows are the network's denoising attempt: each arrow starts at a noisy point and ends at the clean point the network believes it came from.
Press Retrain and watch the loss curve grow from right to left as the epochs tick past. At the start the network is guessing and the error is close to what you get from predicting zero. By the end the arrows point convincingly at the clusters. Drag the slider while the network is fixed and you will see it get worse — it was trained for one noise level, and a different one is out of distribution until you retrain.
Top: the toy data in grey, the noisy points at the selected level in blue, and the network's denoised estimate as pink arrows. Bottom: the training loss, one point per epoch.
The network is seeded, so the same slider value and the same retrain count always produce the same curve. Nothing here uses wall-clock randomness.
A full run is 300 epochs over 160 points — small enough to finish while you read this sentence, large enough to visibly learn.
Epsilon, x₀ and v
Three ways to say the same thing
Given $x_t$ and $t$, the quantities $\varepsilon$, $x_0$ and $v$ are algebraically interchangeable. The forward equation is one linear relation between them, so once you know any one of the three you can recover the others exactly:
- $\varepsilon$-prediction outputs the noise: $\varepsilon_\theta(x_t,t) \approx \varepsilon$. The target is standard normal at every timestep, which is the reason it is the default in most implementations.
- $x_0$-prediction outputs the clean point directly: $x_0 = (x_t - \sqrt{1-\bar\alpha_t}\,\varepsilon)/\sqrt{\bar\alpha_t}$. The target is data, which sounds natural but has a variance that changes with $t$ and, near $t=T$, an enormous amplification factor $1/\sqrt{\bar\alpha_t}$.
- $v$-prediction outputs a rotation-like combination, $v = \sqrt{\bar\alpha_t}\,\varepsilon - \sqrt{1-\bar\alpha_t}\,x_0$. It is the point on the angle bisector: when the signal is strong it looks like noise, when the signal is weak it looks like data.
The demo's network predicts $\varepsilon$, which is why the arrows are computed by first converting its output to $x_0$. You can see that conversion in the readout: the same network output, expressed three ways.
| Target | Formula | Target variance across $t$ |
|---|---|---|
| Noise $\varepsilon$ | $(x_t - \sqrt{\bar\alpha_t}x_0)/\sqrt{1-\bar\alpha_t}$ | Constant: unit normal at every step |
| Clean $x_0$ | $(x_t - \sqrt{1-\bar\alpha_t}\,\varepsilon)/\sqrt{\bar\alpha_t}$ | Data variance, blown up as $\bar\alpha_t\to0$ |
| Velocity $v$ | $\sqrt{\bar\alpha_t}\,\varepsilon - \sqrt{1-\bar\alpha_t}\,x_0$ | Roughly constant across the whole schedule |
Why the parameterisation matters
Same information, very different gradients
If the three targets were exactly recoverable from one another, the choice would be cosmetic. It is not, because a network is a finite approximation trained with finite precision, and the loss weights the different error scales very differently. Training on $x_0$ near the end of the schedule means multiplying the target by $1/\sqrt{\bar\alpha_t}$, which can be a hundredfold amplification of the label: a tiny absolute error becomes a huge loss, gradients explode, and the model spends its capacity on timesteps where the clean image is unrecoverable anyway.
$\varepsilon$-prediction sidesteps that by keeping the target at unit variance everywhere, at the cost of the opposite pathology at high SNR, where the noise is nearly invisible and the model's output is dominated by the prior mean. $v$-prediction splits the difference and was introduced precisely to keep the target's scale stable at both ends of the schedule; it is the default in several of the systems that produce images today.
All three can be made to work in the same architecture, and they often are: a single network can emit all three heads, and a sampler can pick whichever is numerically convenient. What changes is not what the network knows, but how hard it is to learn it. This is a general lesson about diffusion that repeats everywhere — the mathematics is elegant and the numerics decide what actually trains.
Where this shows up
Denoising is gradient estimation
Gradient descent with a schedule
The training loop here is plain stochastic gradient descent with a hand-written backward pass. The nonlinear optimisation guide covers why the learning rate and the loss scale interact the way they do — the same reason $\varepsilon$-prediction is easier to train than $x_0$-prediction.
Regression to a conditional mean
Minimising squared error makes the network converge to $\mathbb{E}[\varepsilon\mid x_t]$, the posterior mean. The linear regression chapter develops the same fact for ordinary least squares: the cheapest thing to learn is always an expectation.
Once a denoiser exists, sampling is a loop: start from noise, ask the network for its noise estimate, subtract a fraction of it, and repeat. The next part runs that loop two ways — one ancestral and stochastic, one deterministic — and counts what each costs.
Further reading
The denoising objective, and the parameterisation debate around it, are the subject of a small literature that is worth reading in order. The DDPM paper defines the loss; the v-prediction paper diagnoses its failure modes; and the progressive-distillation work is the first place the sampler's step count becomes a design variable.
If you read one thing, read the v-prediction paper's analysis of signal-to-noise truncation, because it explains why the same model can look excellent on one benchmark and broken on another.
- Jonathan Ho, Ajay Jain and Pieter Abbeel, "Denoising Diffusion Probabilistic Models", 2020 — the simplified $\varepsilon$-prediction loss.
- Tim Salimans and Jonathan Ho, "Progressive Distillation for Fast Sampling of Diffusion Models", 2022 — introduces $v$-prediction and the signal-to-noise truncation argument.
- Diederik Kingma, Tim Salimans, Ben Poole and Jonathan Ho, "Variational Diffusion Models", 2021 — the continuous-time objective and learned schedules.
- Yang Song and Stefano Ermon, "Generative Modeling by Estimating Gradients of the Data Distribution", 2019 — denoising as score estimation, the bridge to the last part.
- Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning, chapter 6 — backpropagation and the multilayer perceptron, the network trained here.
Cheat sheet
| Term | Meaning here |
|---|---|
| Denoising loss | $\mathbb{E}\lVert\varepsilon-\varepsilon_\theta(x_t,t)\rVert^2$; ordinary regression |
| $\varepsilon$-prediction | Predict the noise; unit-variance target at every $t$ |
| $x_0$-prediction | Predict the clean point; target scale blows up near $t=T$ |
| $v$-prediction | Predict $\sqrt{\bar\alpha_t}\varepsilon-\sqrt{1-\bar\alpha_t}x_0$; stable scale |
| Conditioning on $t$ | One network serves all noise levels via a timestep embedding |
| MLP | The toy network here: 2→16→16→2, tanh, linear output |
| Loss floor | Set by how much of $\varepsilon$ is recoverable from $x_t$ at that $t$ |