What a generative model actually is
Everything in this guide — diffusion, flow matching, latent models, image and audio synthesis — is built on one idea that is easy to say and surprisingly hard to feel: a generative model is a probability distribution you never get to see, and sampling from it means producing a brand-new point that the data never contained. This part makes that idea concrete with a dataset you draw yourself, a model fitted to it in front of you, and a single slider that walks the line between copying the data and learning the shape of the data. Nothing here needs a neural network yet; the next four parts add the machinery that makes the same picture work in dimensions where you cannot draw it by hand.
A distribution you never see
There is a density behind the data, and the data do not contain it
Suppose the world can produce an image, a sentence, or a song. Whatever the medium, there is a set of possible outputs and some notion of how likely each one is — a distribution $p(x)$ over everything the process could make. A photograph of a specific cat on a specific sofa is not stored anywhere; it is one point that a generator, if it were the right one, would assign some density to. In the probability volume the distribution is handed to you. In generative modelling it is the thing you are trying to reconstruct from examples, and you never see it directly.
A dataset is a finite set of samples from $p(x)$. It is a handful of points, not the distribution. This is the whole difficulty in one sentence: a dataset is sparse, and a distribution is everywhere. Two datasets of a hundred images can differ completely while coming from the same distribution, and a model that merely stores the hundred images has learned almost nothing about the process that made them.
The goal of a generative model is therefore not to compress the dataset. It is to fit a distribution $p_\theta(x)$ close enough to the true $p(x)$ that when you draw new points from $p_\theta$, they are plausible in the same way that fresh data would be — sometimes close to a training point, sometimes in the gaps between them, but always in places where the true process could plausibly have put a point.
Hold onto one image before the demo: data are dots, the distribution is the terrain the dots sit on, and modelling means recovering the terrain rather than the dots.
A dataset you draw
Click on the plane to place points
Below is a two-dimensional plane, which is the smallest space in which a distribution is still something you can look at. Every click adds a data point at that location. Two buttons drop in a preset shape — two clusters or a ring — and Undo and Clear let you sculpt a dataset by hand. The grey cloud starts as two clusters because that is the easiest shape to reason about; try the ring afterwards, because a ring has a hole in the middle, and the hole is where the interesting question lives.
Once you have a dataset, ask yourself a question that has no obvious answer: where is the distribution? Not the points — the process that would produce points like these. If I asked you to draw three more points that belong, where would you put them? That judgment is a model, and it is the judgment we are about to hand to an algorithm.
Click anywhere in the frame to add a point. The dataset is whatever is on screen — there is no hidden truth here, only the points you place.
The two-cluster preset is the default because a mixture of two blobs is the simplest distribution whose shape is not just "a blob".
Two properties of a dataset matter everywhere in this guide. The first is its spacing: how far apart neighbouring points are. Spacing is a proxy for how much of the plane the data actually cover, and it sets the natural scale at which a model should invent new points. The second is its shape: a cluster, a ring, a curve, several separate modes. A model that captures the shape but not the spacing overfits; one that captures the spacing but not the shape is too smooth. The demo in the next step uses both.
Sampling from the model
New points, not the ones you drew
Here is the simplest honest model of the dataset you just made: put a small bump of probability at every data point, then draw new points from the union of those bumps. This is a kernel density estimate — the histogram's smoother cousin — and it is a real generative model, fitted to your points, from which we can sample. Each sample picks one data point at random, then wanders a short distance from it. Press Draw 240 samples and the blue points appear over the grey data.
The key readout is novelty: the typical distance from a model sample to the nearest data point, measured in units of the dataset's own spacing. If that number is zero, the model has learned to copy. If it is comfortably above zero, the model is putting points where the data are not but easily could have been.
Grey: your data. Blue: 240 fresh samples from the fitted kernel-density model. No sample is a data point unless the jitter is zero.
Sampling is random, but this page is seeded: the same slider value always produces the same 240 points, so you can compare settings fairly.
Notice what sampling buys you. The model has a compact description — the dataset plus one bandwidth — and from that it can emit as many points as you like. The samples are not memorised outputs; they are draws from a distribution that was fitted to the data. That is the entire contract of a generative model, and the rest of this guide is about doing it in spaces where "put a bump at every point" is hopeless: a million-dimensional image, where a bump per training example is not a model, it is a lookup table.
Memorising versus modelling
The jitter slider is the difference, and it is one number
Drag the jitter slider to zero and press the button again. Every sample lands exactly on a data point. The model has become a copy machine: it can only ever return one of the inputs, so it cannot produce anything the dataset did not already contain. Now drag the slider up. The samples spread out, following the shape of the data — the clusters stay clustered, the ring stays a ring — and the novelty readout climbs. Some samples fall in the gaps between points. That is generalisation, made visible.
Push it too far and the model breaks in the opposite direction: the bumps swell until clusters merge, the hole in the ring fills in, and the samples stop looking like the data at all. So there is a sweet spot, and finding it is what every training algorithm in this guide does — not by one slider, but by fitting millions of parameters so that the implied distribution matches the data without collapsing onto it. In diffusion the equivalent knob is the noise schedule and the amount of training; too little and the model memorises, too much and it over-smooths.
Novelty against jitter, computed from the current dataset. The dashed line is where the slider is right now.
At jitter zero the curve sits exactly on the axis: every sample is a training point. It rises steeply at first — the first bit of jitter is what breaks the copy machine — and then flattens, because past a certain spread the nearest data point is no closer than the spacing between arbitrary points in the plane.
Change the dataset on the previous canvas and this curve is recomputed: a ring and a pair of clusters have different spacing, so the same jitter means different things.
The same tension is the reason this guide spends four more parts on noise. Adding noise to an image is a way of moving it off the data manifold into that between-space, and training a network to remove the noise is a way of forcing it to model the between-space rather than the points. The forward process is next.
Where this shows up
The same picture, in higher dimensions
Densities and conditionals
The probability guide supplies the language this part uses informally: densities, mixtures, and the conditional Gaussian that the forward process will lean on. The probabilistic machine-learning chapter walks the same "fit a distribution to points" picture and ends at the one-page diffusion on-ramp.
Reading as well as making
The sibling volume Multimodal Models, Interactively takes the same distributions and conditions them on text, images and audio. The joint distribution you sketch here becomes a conditional one there: instead of $p(x)$, a model learns $p(x \mid \text{prompt})$.
In the next part the points stop being two-dimensional dots and become pixels. A 64×64 image is a point in 12,288 dimensions, which is far too many to plot but exactly the same idea — a dataset of points, a distribution behind them, and a model that must learn the terrain rather than the dots. The machinery that makes it work is a procedure that destroys a data point with noise, one small step at a time, until nothing but the distribution remains.
Further reading
The references below are the canonical treatments of generative models as density estimation, in increasing order of machinery. If you take one idea away, take the separation between a dataset and the distribution behind it; almost every confusion later in this volume comes from collapsing those two.
Bishop's chapter on density estimation is the cleanest introduction to kernels and mixtures; Goodfellow's tutorial and Kingma and Welling's VAE paper are where the modern framing of "sample from a learned distribution" begins; and the DDPM paper is the point at which this guide's next four parts start. 3Blue1Brown's videos are the visual reference for the style of the demos here.
- Christopher Bishop, Pattern Recognition and Machine Learning, chapter 2 — binary and multinomial distributions, the Gaussian, and kernel density estimation.
- Ian Goodfellow, "NIPS 2016 Tutorial: Generative Adversarial Networks" — the taxonomy of generative models as density estimation, explicit and implicit.
- Diederik Kingma and Max Welling, "Auto-Encoding Variational Bayes", 2013 — the latent-variable view of a generative model and the reparameterisation trick.
- Jonathan Ho, Ajay Jain and Pieter Abbeel, "Denoising Diffusion Probabilistic Models", 2020 — the model the next part begins to take apart.
- 3Blue1Brown, "But what is a diffusion model?" — the visual intuition for adding and removing noise.
Cheat sheet
| Term | Meaning here |
|---|---|
| Distribution $p(x)$ | The process that generates data; never observed, only inferred |
| Dataset | A finite set of samples; sparse, and not the distribution |
| Generative model | A fitted distribution $p_\theta(x)$ you can sample from |
| Sample | One draw from the model; ideally novel, not a copied input |
| Memorisation | Outputs confined to the training points; zero novelty |
| Generalisation | Outputs that respect the shape of the data but fill the gaps |
| Kernel density estimate | A bump at every data point; the simplest real generative model |
| Jitter / temperature | The width of those bumps; the dial between copying and inventing |