Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A distribution you never see

There is a density behind the data, and the data do not contain it

Suppose the world can produce an image, a sentence, or a song. Whatever the medium, there is a set of possible outputs and some notion of how likely each one is — a distribution $p(x)$ over everything the process could make. A photograph of a specific cat on a specific sofa is not stored anywhere; it is one point that a generator, if it were the right one, would assign some density to. In the probability volume the distribution is handed to you. In generative modelling it is the thing you are trying to reconstruct from examples, and you never see it directly.

A dataset is a finite set of samples from $p(x)$. It is a handful of points, not the distribution. This is the whole difficulty in one sentence: a dataset is sparse, and a distribution is everywhere. Two datasets of a hundred images can differ completely while coming from the same distribution, and a model that merely stores the hundred images has learned almost nothing about the process that made them.

The goal of a generative model is therefore not to compress the dataset. It is to fit a distribution $p_\theta(x)$ close enough to the true $p(x)$ that when you draw new points from $p_\theta$, they are plausible in the same way that fresh data would be — sometimes close to a training point, sometimes in the gaps between them, but always in places where the true process could plausibly have put a point.

💡 By the end of this part you'll be able to state what a generative model is as a distribution you never see, build a dataset by hand, sample from a fitted version of it, and explain why a model that can only reproduce its training data has failed even when every lookup is exactly right.

Hold onto one image before the demo: data are dots, the distribution is the terrain the dots sit on, and modelling means recovering the terrain rather than the dots.

2

A dataset you draw

Click on the plane to place points

Below is a two-dimensional plane, which is the smallest space in which a distribution is still something you can look at. Every click adds a data point at that location. Two buttons drop in a preset shape — two clusters or a ring — and Undo and Clear let you sculpt a dataset by hand. The grey cloud starts as two clusters because that is the easiest shape to reason about; try the ring afterwards, because a ring has a hole in the middle, and the hole is where the interesting question lives.

Once you have a dataset, ask yourself a question that has no obvious answer: where is the distribution? Not the points — the process that would produce points like these. If I asked you to draw three more points that belong, where would you put them? That judgment is a model, and it is the judgment we are about to hand to an algorithm.

Click anywhere in the frame to add a point. The dataset is whatever is on screen — there is no hidden truth here, only the points you place.

The two-cluster preset is the default because a mixture of two blobs is the simplest distribution whose shape is not just "a blob".

Two properties of a dataset matter everywhere in this guide. The first is its spacing: how far apart neighbouring points are. Spacing is a proxy for how much of the plane the data actually cover, and it sets the natural scale at which a model should invent new points. The second is its shape: a cluster, a ring, a curve, several separate modes. A model that captures the shape but not the spacing overfits; one that captures the spacing but not the shape is too smooth. The demo in the next step uses both.

3

Sampling from the model

New points, not the ones you drew

Here is the simplest honest model of the dataset you just made: put a small bump of probability at every data point, then draw new points from the union of those bumps. This is a kernel density estimate — the histogram's smoother cousin — and it is a real generative model, fitted to your points, from which we can sample. Each sample picks one data point at random, then wanders a short distance from it. Press Draw 240 samples and the blue points appear over the grey data.

The key readout is novelty: the typical distance from a model sample to the nearest data point, measured in units of the dataset's own spacing. If that number is zero, the model has learned to copy. If it is comfortably above zero, the model is putting points where the data are not but easily could have been.

Grey: your data. Blue: 240 fresh samples from the fitted kernel-density model. No sample is a data point unless the jitter is zero.

Sampling is random, but this page is seeded: the same slider value always produces the same 240 points, so you can compare settings fairly.

Notice what sampling buys you. The model has a compact description — the dataset plus one bandwidth — and from that it can emit as many points as you like. The samples are not memorised outputs; they are draws from a distribution that was fitted to the data. That is the entire contract of a generative model, and the rest of this guide is about doing it in spaces where "put a bump at every point" is hopeless: a million-dimensional image, where a bump per training example is not a model, it is a lookup table.

4

Memorising versus modelling

The jitter slider is the difference, and it is one number

Drag the jitter slider to zero and press the button again. Every sample lands exactly on a data point. The model has become a copy machine: it can only ever return one of the inputs, so it cannot produce anything the dataset did not already contain. Now drag the slider up. The samples spread out, following the shape of the data — the clusters stay clustered, the ring stays a ring — and the novelty readout climbs. Some samples fall in the gaps between points. That is generalisation, made visible.

Push it too far and the model breaks in the opposite direction: the bumps swell until clusters merge, the hole in the ring fills in, and the samples stop looking like the data at all. So there is a sweet spot, and finding it is what every training algorithm in this guide does — not by one slider, but by fitting millions of parameters so that the implied distribution matches the data without collapsing onto it. In diffusion the equivalent knob is the noise schedule and the amount of training; too little and the model memorises, too much and it over-smooths.

Novelty against jitter, computed from the current dataset. The dashed line is where the slider is right now.

At jitter zero the curve sits exactly on the axis: every sample is a training point. It rises steeply at first — the first bit of jitter is what breaks the copy machine — and then flattens, because past a certain spread the nearest data point is no closer than the spacing between arbitrary points in the plane.

Change the dataset on the previous canvas and this curve is recomputed: a ring and a pair of clusters have different spacing, so the same jitter means different things.

⚠ Exact outputs are not the goal. A model that reproduces its training set perfectly would score zero on any honest test of novelty, yet it is trivially easy to build — it is the dataset itself. Everything that makes a generative model useful happens in the space between the training points, and that space is exactly where there is no data to check against.

The same tension is the reason this guide spends four more parts on noise. Adding noise to an image is a way of moving it off the data manifold into that between-space, and training a network to remove the noise is a way of forcing it to model the between-space rather than the points. The forward process is next.

5

Where this shows up

The same picture, in higher dimensions

Probability

Densities and conditionals

The probability guide supplies the language this part uses informally: densities, mixtures, and the conditional Gaussian that the forward process will lean on. The probabilistic machine-learning chapter walks the same "fit a distribution to points" picture and ends at the one-page diffusion on-ramp.

Multimodal

Reading as well as making

The sibling volume Multimodal Models, Interactively takes the same distributions and conditions them on text, images and audio. The joint distribution you sketch here becomes a conditional one there: instead of $p(x)$, a model learns $p(x \mid \text{prompt})$.

In the next part the points stop being two-dimensional dots and become pixels. A 64×64 image is a point in 12,288 dimensions, which is far too many to plot but exactly the same idea — a dataset of points, a distribution behind them, and a model that must learn the terrain rather than the dots. The machinery that makes it work is a procedure that destroys a data point with noise, one small step at a time, until nothing but the distribution remains.

Further reading

The references below are the canonical treatments of generative models as density estimation, in increasing order of machinery. If you take one idea away, take the separation between a dataset and the distribution behind it; almost every confusion later in this volume comes from collapsing those two.

Bishop's chapter on density estimation is the cleanest introduction to kernels and mixtures; Goodfellow's tutorial and Kingma and Welling's VAE paper are where the modern framing of "sample from a learned distribution" begins; and the DDPM paper is the point at which this guide's next four parts start. 3Blue1Brown's videos are the visual reference for the style of the demos here.

Cheat sheet

TermMeaning here
Distribution $p(x)$The process that generates data; never observed, only inferred
DatasetA finite set of samples; sparse, and not the distribution
Generative modelA fitted distribution $p_\theta(x)$ you can sample from
SampleOne draw from the model; ideally novel, not a copied input
MemorisationOutputs confined to the training points; zero novelty
GeneralisationOutputs that respect the shape of the data but fill the gaps
Kernel density estimateA bump at every data point; the simplest real generative model
Jitter / temperatureThe width of those bumps; the dial between copying and inventing
6

Check your understanding

0/4 answered