Destroying an image with noise
Diffusion models are trained on a job that sounds absurd: look at a picture, add noise to it, and ask a network to name the noise. The surprise is that this single trick turns an impossible density-estimation problem into a sequence of easy regressions. This part builds the first half of that trick — the forward process that turns a clean image into a field of static — and shows the closed-form shortcut that makes training cheap. By the end you will be able to read any noise schedule off a graph and say exactly how much signal is left at any step.
What the forward process does
A small amount of noise, repeated many times
Start with a clean image $x_0$. Add a little Gaussian noise and you get $x_1$. Add a little more and you get $x_2$, and so on for a thousand steps until $x_T$ is indistinguishable from pure static. This chain is the forward process, and it is not learned — it is a fixed recipe. Each step is $x_t = \sqrt{1-\beta_t}\,x_{t-1} + \sqrt{\beta_t}\,\varepsilon_t$ with $\varepsilon_t \sim \mathcal{N}(0, I)$, where the coefficients are chosen so that the signal shrinks and the noise grows by matching amounts.
Two design choices are worth noticing now, because they are the reason everything later simplifies. The noise is standard Gaussian, so the whole chain stays in the family of Gaussian distributions; and the scaling by $\sqrt{1-\beta_t}$ keeps the variance of $x_t$ at roughly one, so the numbers never blow up or vanish. That second choice is the difference between diffusion and a random walk that merely gets larger.
Slide the timestep below and watch the blobs dissolve. The patch grid is drawn on top because a diffusion network does not see the image as pixels — it sees it as a grid of patches, and the noise is applied to all of them at once.
A procedurally generated 64×64 image at timestep $t$, with an 8×8 patch grid overlaid. The same noise field is reused at every $t$, so dragging the slider shows one image being destroyed, not a new image each time.
Switch schedules at the same $t$ and compare: the cosine curve is still recognisable where the linear one is already gone.
The closed form
Jump to any timestep in one line
The step-by-step chain is a definition, but training does not want to run it. If you fold the recursion, the sum of many small Gaussians is again a Gaussian, and you get the single most useful equation in the whole subject:
$$ q(x_t \mid x_0) = \mathcal{N}\!\left(\sqrt{\bar\alpha_t}\,x_0,\; (1-\bar\alpha_t) I\right) $$
where $\bar\alpha_t = \prod_{s=1}^{t}(1-\beta_s)$ is the cumulative product of the signal coefficients. In words: to produce $x_t$ you do not need steps $1$ through $t-1$; you just scale the clean image by $\sqrt{\bar\alpha_t}$, add Gaussian noise of standard deviation $\sqrt{1-\bar\alpha_t}$, and you are done. The mean is a shrunken copy of $x_0$ and the variance is the noise that has accumulated, and the two always sum to one.
The lower canvas shows why this matters. The grey histogram is the distribution of pixel values in the clean image; the pink one is the same pixels after the forward process at the current $t$. At small $t$ the pink bars sit on top of the grey ones, just a little wider. As $t$ grows the pink distribution flattens toward a standard normal, which is exactly the statement that $q(x_t \mid x_0)$ has forgotten $x_0$.
Pixel-value distributions: original ($x_0$) against noisy ($x_t$) at the timestep chosen above.
The forward process is a bridge between two distributions that both have a simple name. At $t=0$ it is a point mass at your data. At large $t$ it is standard normal noise. Everything a diffusion model does happens in between, and the closed form lets it sample any point on that bridge for free.
This is also why training is cheap: at every gradient step you draw a random $t$, jump straight to $x_t$, and regress. No rollout, no backpropagation through time.
ᾱ, SNR and the schedules
Two curves that describe an entire chain
Everything you need to know about a noise schedule is in one function of time, $\bar\alpha_t$, and everything about how hard the denoising problem is at time $t$ is in the signal-to-noise ratio $\mathrm{SNR}(t) = \bar\alpha_t / (1-\bar\alpha_t)$. High SNR means the image is mostly intact and the model has an easy job; low SNR means the image is mostly noise and the model is guessing at a coarse scale. The crossover where SNR equals one is where the picture stops being visible by eye, and it lands at very different times for different schedules.
The first plot traces $\bar\alpha_t$ for the linear schedule of the original DDPM paper and the cosine schedule of Nichol and Dhariwal. The second plots $\log_{10}\mathrm{SNR}$, which straightens out the exponential decay and makes the difference in slope the whole story. The dashed line marks the timestep selected above, so you can read off exactly where the demo currently is.
Top: $\bar\alpha_t$ for both schedules. Bottom: $\log_{10}\mathrm{SNR}(t)$. The vertical dashed line is the current timestep.
The linear schedule — $\beta$ from $10^{-4}$ to $0.02$ — spends the first half of its budget barely touching the image and then destroys it quickly near the end. The cosine schedule was designed to fix exactly that imbalance: it removes signal more evenly, so the middle of training is not wasted on timesteps where nothing interesting is happening.
Both curves are computed with accent for linear and accent-2 for cosine, matching the two buttons in the demo.
What the schedule choice changes
A schedule is a curriculum for the denoiser
A schedule sets the distribution of difficulty the network sees during training. Timesteps where $\mathrm{SNR}$ is enormous are almost trivial — the answer is nearly visible — and timesteps where it is tiny are nearly impossible, because the image has been fully forgotten. The interesting range is the band around moderate SNR, and a schedule that spends most of its steps there gives the network more useful gradients per unit of compute.
That is the real reason cosine tends to beat linear at the same number of sampling steps. It is not that one is more mathematically correct; both define a valid $q$, and the reverse process can in principle invert either. It is that cosine puts more of its resolution where the denoising problem is hard, so a sampler with a fixed step budget spends its steps wisely.
The choice interacts with everything downstream. More steps at high SNR sharpen detail; more steps near the end fix global composition, because at low SNR the model can only move large-scale structure. Later parts will pick a sampler and a step count, and those choices are usually made by looking at the SNR curve and spending steps where it is changing fastest.
One more quantity is worth naming before we leave: at any $t$, the quantity $\sqrt{1-\bar\alpha_t}$ is the standard deviation of the noise that has been added, and it is exactly what the network will be asked to predict. The denoiser's target is the noise, and the score function is that same prediction up to a known factor — but that is the subject of the fifth part.
Where this shows up
Forward processes well beyond images
Gaussians compose
The closed form is just the fact that a sum of independent Gaussians is Gaussian, and the scaling keeps the result in the same family. The probability guide develops the algebra; the probabilistic ML chapter uses it to derive the variance-preserving forward chain from scratch.
Noise is modality-agnostic
Nothing in $q(x_t \mid x_0)$ cares whether $x_0$ is an image, a spectrogram or a video clip. The same closed form runs on all of them, which is why the later parts of this volume reuse this exact machinery in sound and video.
If you understand that every timestep is one Gaussian sample away from the clean data, you understand the first half of diffusion. The second half is the interesting half: given $x_t$, estimate the noise that was added, and then use that estimate to walk backwards. The next part trains a real denoiser in the browser and watches its loss fall.
Further reading
The forward process appears in every diffusion paper, usually without much ceremony, because it is the easy half. These references give it the care it deserves, and the schedules they introduce are the ones used in production systems today.
Sohl-Dickstein and coauthors introduced the non-equilibrium thermodynamics framing; Ho, Jain and Abbeel turned it into the modern denoising objective; Nichol and Dhariwal fixed the schedule; and Song and colleagues show how far the idea generalises. If you read one section, read the schedule comparison in Nichol and Dhariwal.
- Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan and Surya Ganguli, "Deep Unsupervised Learning using Nonequilibrium Thermodynamics", 2015 — the origin of the forward and reverse process.
- Jonathan Ho, Ajay Jain and Pieter Abbeel, "Denoising Diffusion Probabilistic Models", 2020 — the linear schedule and the simplified training loss.
- Alexander Nichol and Prafulla Dhariwal, "Improved Denoising Diffusion Probabilistic Models", 2021 — the cosine schedule and the argument for spending steps where SNR is changing.
- Yang Song, Jascha Sohl-Dickstein and coauthors, "Score-Based Generative Modeling through Stochastic Differential Equations", 2020 — the continuous-time view that makes schedules a matter of choosing a variance curve.
Cheat sheet
| Symbol | Meaning here |
|---|---|
| $\beta_t$ | Noise added at step $t$; small, and fixed in advance |
| $\alpha_t = 1-\beta_t$ | Signal kept at step $t$ |
| $\bar\alpha_t = \prod_{s\le t}\alpha_s$ | Cumulative signal after $t$ steps |
| $q(x_t\mid x_0)$ | $\mathcal{N}(\sqrt{\bar\alpha_t}\,x_0,\,(1-\bar\alpha_t)I)$ — the closed form |
| $\sqrt{1-\bar\alpha_t}$ | Standard deviation of the added noise; the denoiser's target |
| SNR$(t)$ | $\bar\alpha_t/(1-\bar\alpha_t)$; how easy denoising is at $t$ |
| Linear schedule | $\beta$ ramps $10^{-4}\to0.02$; destroys late and fast |
| Cosine schedule | Removes signal evenly; more useful steps in the middle |