Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Glossary

Every term, linked to the part that introduces it

💡 Filter the list. Type any fragment — a term, a metric, a model, a technology — and the table hides non-matching rows and reports how many remain. Matching is case-insensitive and looks at both the term and its definition.
TermWhat it meansIntroduced in
A
AdaLN-ZeroConditioning that lets a DiT block stay a plain transformer: the timestep and text embeddings produce scale, shift and gate values, and the modulating projections are initialised to zero so each block starts as the identity.Part 9
AliasingA frequency above the Nyquist limit folding back and appearing as a lower one; the artefact that makes a low-pass filter mandatory before any downsampling.Part 16
Ancestral samplerA sampler that injects fresh noise at every step, so the trajectory is stochastic; the original DDPM reverse process is its canonical form.Part 4
Attention (cross against joint)Two ways to inject conditioning: a separate cross-attention pass against the prompt embedding, or text and image tokens concatenated into one sequence that self-attends jointly.Part 10
B
BackboneThe network that predicts the noise, the clean image or the velocity: a U-Net with a resolution pyramid, or a transformer over patched latents.Part 8
Barge-inA user interrupting a spoken reply. A duplex system has to detect it and stop speaking inside the interaction budget, which is a latency problem rather than an accuracy one.Part 19
BitrateBits per second a codec emits: frame rate times codebooks times bits per codebook, the number that makes a codec a token budget rather than an audio format.Part 17
C
C2PAContent Credentials: a cryptographically signed manifest attached to a file that records who made it, with which tool and what edits. Strong while it is attached, and gone the moment a platform strips metadata.Part 20
Causal 3-D VAEA video compressor whose temporal convolutions see only the current and past frames, so no latent frame can depend on the future.Part 14
Classifier-free guidance (CFG)Training with the condition dropped at random, then extrapolating between the conditional and unconditional predictions at sampling time; the guidance scale is the extrapolation factor.Part 10
CLIPThe contrastive text-image encoder used both as the text encoder in the early stack and as the scoring model behind CLIPScore.Part 10
CLIPScoreThe cosine between a CLIP image embedding and a text embedding, used as a reference-free alignment metric. It measures prompt adherence and is blind to realism, anatomy and aesthetics.Part 20
CodebookThe learned vector table a quantiser snaps its input to. Residual quantisation stacks several codebooks so each stage corrects what the previous one left behind.Part 17
Consistency distillationTraining the model to map any point on a sampling trajectory straight to its endpoint, which is what makes one- and few-step generation possible.Part 12
ControlNetA trainable copy of the backbone driven by a conditioning image such as edges, depth or pose, injected back into the frozen model's blocks.Part 13
CTCConnectionist temporal classification: an alignment-free output layer that permits a blank symbol and collapses repeats, producing monotonic text from a frame sequence.Part 19
D
DDIMA deterministic, non-ancestral sampler built from the same trained model; it removes the per-step noise injection and allows larger jumps between timesteps.Part 4
DDPMThe denoising diffusion probabilistic model: a fixed forward noising chain plus a learned reverse chain trained on the variational bound.Part 4
DiffusionThe family of generative models that destroy data with noise and learn to reverse the destruction one step at a time.Part 1
DiTA diffusion transformer: the latent is patched into tokens and processed by plain transformer blocks conditioned through adaptive layer norm.Part 9
Distillation (diffusion)Training a few-step student from a many-step teacher, by progressive halving, consistency objectives, adversarial losses or shortcut models.Part 12
DPM-Solver++A multistep exponential-integrator solver for the probability-flow ODE; the usual default when a strong epsilon- or x0-prediction model has to run in a few dozen steps.Part 11
DuplexFull-duplex speech: listening while speaking, which turns turn-taking and interruption into latency budgets rather than transcription problems.Part 19
E
EloA rating scale fitted to pairwise votes, usually through a Bradley-Terry model. Relative, prompt-set dependent, and moved by voter style preference.Part 20
Epsilon-predictionTraining the denoiser to output the noise that was added rather than the clean image; the original diffusion parameterisation and still the most common.Part 3
Euler solverThe simplest ODE integrator, one function evaluation per step; the k-diffusion and flow-matching default because straight paths make a first-order method accurate.Part 11
F
FIDFrechet Inception Distance: the Frechet distance between Gaussians fitted to real and generated feature clouds. It measures marginal realism, cannot see the conditioning, and is biased upward at small sample counts.Part 20
Flow matchingTraining a velocity field that carries a simple distribution to the data along a chosen path, usually a straight line between noise and sample.Part 6
Frame rate (codec)Codec frames per second, around fifty to seventy-five in modern audio codecs; one of the two factors of the token rate an audio language model sees.Part 17
FVDFrechet Video Distance: FID computed on video features, so spatial and temporal quality collapse into a single scalar with wide error bars.Part 20
G
Gaussian posteriorThe closed-form conditional that makes the forward process reversible in one step: given a noisy image and the clean one, the previous state is Gaussian.Part 2
GenAI-ArenaA human pairwise-vote arena for generative models whose leaderboard is an Elo fit over the votes that users cast on submitted prompts.Part 20
Guidance scaleThe weight on the conditional minus unconditional difference; raising it improves prompt adherence and costs diversity and can oversaturate the image.Part 10
H
Heun solverA second-order predictor-corrector ODE solver: one extra function evaluation per step buys substantially lower discretisation error.Part 11
Human Preference Score (HPS)A preference model trained on the HPD v2 dataset of human choices, used as a differentiable proxy for taste. It inherits the annotator pool and is gamed once optimised against.Part 20
I
InpaintingRegenerating a masked region while keeping the rest of the image, done by mixing noise into the latent inside the mask at every step and keeping the known region consistent.Part 13
IP-AdapterA small adapter that injects a reference image as an extra conditioning signal into a frozen backbone, so a photograph can steer style or subject.Part 13
K
Keyframe conditioningSupplying first, last or intermediate frames to a video model so a clip starts and ends where the user wants it to.Part 15
L
Langevin dynamicsSampling by following the score field with a noisy gradient step; the ancestral sampler is one discretisation of it.Part 5
LatentThe compressed representation a VAE produces, about eight times smaller along each spatial axis, in which diffusion actually runs.Part 7
LoRAA low-rank update to frozen weights; the standard way to personalise or steer a model without fine-tuning all of it.Part 13
M
Mel scaleA perceptual frequency scale that spaces bins by how they sound rather than by hertz; the vertical axis of the spectrogram most audio models consume.Part 16
MemorisationA model reproducing a training example rather than generating a new one, concentrated in duplicated and atypical images and measured by nearest-neighbour similarity.Part 20
N
Noise scheduleThe sequence of betas that sets how fast signal is destroyed, summarised by its cumulative product alpha-bar.Part 2
Nyquist frequencyHalf the sample rate: the highest frequency a given rate can represent without aliasing.Part 16
O
ODE samplerA sampler that integrates the probability-flow ODE deterministically, in contrast to the stochastic ancestral sampler; the family Euler, Heun and DPM-Solver++ belong to.Part 11
One-step generationProducing an image in a single forward pass, by distillation or by training a new model; the extreme end of the steps-against-quality trade.Part 12
P
PatchifyCutting a latent into fixed tiles and flattening them into a token sequence; the operation that turns an image into transformer input.Part 9
PickScoreA preference model trained on the Pick-a-Pic dataset of user choices. It correlates with human judgement better than CLIPScore and is gamed once it becomes a training signal.Part 20
Progressive distillationHalving a sampler's step count repeatedly, each student learning to reproduce two teacher steps in one.Part 12
ProvenanceEvidence about how a piece of media was made: a watermark carried in the pixels, a signed manifest in the metadata, or both. Absence of either is not evidence of authenticity.Part 20
Q
Quantization (serving)Storing weights and activations at FP8, INT8 or FP4 to cut bytes and raise the arithmetic rate. The accuracy bill has to be measured on images rather than inherited from language-model results.Part 20
R
Rectified flowFlow matching on straight paths plus a reflow step that straightens them further, so a first-order solver is accurate at few steps.Part 6
Residual vector quantisation (RVQ)Quantising the residual of the previous stage across several codebooks, so reconstruction error falls with each added stage.Part 17
RNN-TA transducer that combines a prediction network, an encoder and a joint network to emit tokens while streaming, without CTC's frame-independence assumption.Part 19
S
SamplerThe procedure that walks the trained model from noise to a sample: ancestral, a first-order ODE step, or a higher-order multistep solver.Part 11
SDEditEditing by adding noise to an image up to a chosen level and denoising back, which sets how far the edit is allowed to travel from the original.Part 13
Score functionThe gradient of the log density with respect to the data; the quantity the denoiser implicitly estimates and the object Langevin sampling follows.Part 5
Signal-to-noise ratio (SNR)The ratio of signal variance to noise variance at a timestep; its logarithm is the natural axis for schedules and for solver error.Part 2
Spatiotemporal patchA latent patch spanning several frames; the token unit of a video transformer after the causal 3-D VAE has compressed space and time.Part 14
Step cachingReusing backbone features across adjacent sampler steps and skipping the recomputation where the change is small; the cheapest server-side speedup.Part 20
SynthIDA sampling-time watermark whose key-dependent residual is spread over the whole image and recovered by a detector that holds the key.Part 20
T
Temporal cascadeA base model at low resolution plus a temporal upsampler, so clip length grows without making attention quadratic in the whole clip.Part 15
Text encoderThe frozen language or CLIP model that turns the prompt into the conditioning sequence the backbone attends to, and the reason prompt length costs prefill time.Part 10
Text to speech (TTS)Turning text into audio: alignment or duration prediction first, then acoustic tokens or a spectrogram, then a vocoder.Part 18
Token rateAudio tokens per second: frame rate times codebooks. It is the number that decides whether a language model over audio is tractable at all.Part 17
U
U-NetA convolutional backbone with a resolution pyramid and skip connections; the image-generation workhorse before transformers took over the latent.Part 8
V
VAEThe variational autoencoder whose encoder compresses an image into a latent and whose decoder turns a latent back into pixels.Part 7
v-predictionPredicting a velocity-like combination of the clean image and the noise, which behaves better than epsilon-prediction at the extremes of the schedule.Part 3
Velocity fieldThe field a flow-matching model learns: the direction and speed of the path that carries noise to data.Part 6
Voice activity detection (VAD)Deciding when a person is speaking; the trigger that ends a turn and starts the recogniser in a duplex system.Part 19
Voice cloningReproducing an unseen speaker from a short reference recording: zero-shot with a codec language model, or few-shot with an adapter.Part 18
VocoderThe component that turns a spectrogram or codec tokens back into a waveform, in current systems often a small diffusion or flow model of its own.Part 18
W
WatermarkA signal embedded in the output so it can be identified later, either distributed through the pixels or declared in a signed manifest.Part 20
World modelThe claim that a video generator has learned dynamics that can be controlled and queried rather than merely rendered.Part 15
X
x0-predictionParameterising the denoiser to output the clean image directly: convenient for clipping, but badly scaled at high noise.Part 3
Z
Zero-shot cloningCloning a voice from a reference clip with no per-speaker fine-tuning, which is what codec language models made routine.Part 18
2

Model lineage

From a 512-pixel latent U-Net to a flow-matching transformer

Read the table as a change of objective and of backbone rather than as a list of releases. The first generation diffused a latent with a U-Net and predicted noise; the current generation predicts a velocity along a straight path with a transformer, and the only reason the two can be discussed together is that both are trained by predicting something about a noised input and both are sampled by walking backwards. Where a parameter count or a training resolution is undisclosed, the cell says so rather than guessing.

ModelReleasedResolutionParametersObjectiveBackbone
Stable Diffusion 1.52022-10512 px square0.86 B U-Net, 0.08 B VAE, CLIP ViT-L text encoderDDPM epsilon-prediction, classifier-free guidanceLatent U-Net
Stable Diffusion XL2023-071024 px square2.6 B U-Net, two text encoders (CLIP ViT-L and OpenCLIP ViT-bigG)Epsilon-prediction with size and crop conditioning; base plus refinerLatent U-Net
Stable Diffusion 3.52024-101024 px, up to 1440 px8 B Large, 2.5 B MediumRectified flow with logit-normal timestep samplingMMDiT
FLUX.12024-081024 to 2048 px12 BRectified flow matching; a guidance-distilled schnell variant runs in a few stepsMMDiT, double and single stream blocks
Qwen-Image2025-081024 px and above, arbitrary aspect ratio20 BFlow matching, with strong text renderingMMDiT
Veo 32025-05720p to 1080p video at 24 fps, native audioUndisclosedFlow matching over video and audio latentsDiT with factorised attention
Sora 22025-09Up to 1080p video with synchronised audioUndisclosedSpatiotemporal latent diffusion with flow matchingDiT over spatiotemporal patches

Released dates are public announcement dates; parameter counts are the published configuration at release, and undisclosed means the developer has not published a figure. Text encoders and VAEs are counted separately from the diffusion backbone where the counts were published separately.

3

Sampler comparison

Deterministic or not, what order, and what it costs

The sampler is where the step budget is decided, and the whole table follows from two questions: does the update inject noise, and how many function evaluations does one step use. Ancestral sampling injects noise, so it is stochastic and needs many small steps. The ODE solvers do not, so they can take larger jumps, and a higher-order method spends more model calls per step to make those jumps accurate. Flow-matching models make the straightest paths, which is why a first-order Euler step is competitive for them at step counts that a curved DDPM path could never reach.

SamplerDeterministic?OrderTypical stepsNotes
DDPM ancestralNoFirst250 to 1000Adds a fresh noise draw at every step, matching the training objective. The reference behaviour, and by far the most model calls.
DDIMYes when eta is zeroFirst20 to 100The same trained model with the per-step noise removed; supports skipping timesteps, and the eta knob interpolates back toward ancestral sampling.
DPM-Solver++YesSecond to third15 to 30A multistep exponential integrator in log-SNR. The usual default for epsilon- and x0-prediction models, and the point where a second order pays for itself.
Heun (predictor-corrector)YesSecond25 to 50One extra function evaluation per step to correct the Euler prediction. Used by the EDM formulation and cheap at low model scales.
EulerYesFirst20 to 50 on a curved path, far fewer on a straight oneThe baseline ODE integrator: one model call per step, no memory of previous steps. Its accuracy depends entirely on how straight the path is.
Flow matching, Euler ODEYesFirst20 to 50, or 1 to 4 after distillationThe current frontier default: straight paths make a first-order step accurate, and a guidance-distilled or consistency student can collapse the count to a handful.

Step counts are the ranges reported for the samplers in practice, not hard limits: quality depends on the model, the schedule spacing and the guidance scale. A distilled model changes the row's typical count rather than its order.

4

Formulas to keep

Four expressions the guide builds everything from

The schedule fixes how much signal is left at each timestep, the DDIM step is the deterministic update written in terms of that schedule, guidance is an extrapolation between two predictions of the same model, and the flow-matching loss is what the frontier trainers replaced the noise-prediction loss with. Everything else in the series is a consequence of these.

The noise schedule

$$\bar\alpha_t = \prod_{s=1}^{t}(1-\beta_s), \qquad x_t = \sqrt{\bar\alpha_t}\,x_0 + \sqrt{1-\bar\alpha_t}\,\epsilon$$

Beta is the per-step variance of the forward process and alpha-bar is its cumulative signal retention. The second expression is the one-step forward kernel: a sample at any timestep is a closed-form mixture of the clean image and a standard normal draw, which is what makes training a single-step denoiser possible.

The DDIM step

$$x_{t-1} = \sqrt{\bar\alpha_{t-1}}\,\hat{x}_0 + \sqrt{1-\bar\alpha_{t-1}}\,\epsilon_\theta(x_t,t), \qquad \hat{x}_0 = \frac{x_t - \sqrt{1-\bar\alpha_t}\,\epsilon_\theta(x_t,t)}{\sqrt{\bar\alpha_t}}$$

Estimate the clean image from the model's noise prediction, then re-noise it to the next timestep's level. No noise is added, so the trajectory is deterministic and the same seed and prompt give the same image twice.

Classifier-free guidance

$$\tilde{\epsilon}_\theta(x_t, c) = \epsilon_\theta(x_t, \varnothing) + w\,\bigl(\epsilon_\theta(x_t, c) - \epsilon_\theta(x_t, \varnothing)\bigr)$$

One model is evaluated twice, with and without the condition, and the difference is scaled by the guidance weight w. At w equal to one it is ordinary conditional sampling; above one prompt adherence improves while diversity falls and colours oversaturate.

Flow-matching loss

$$\mathcal{L} = \mathbb{E}_{t,\,x_0,\,x_1}\left\lVert v_\theta(x_t, t) - (x_1 - x_0) \right\rVert^2, \qquad x_t = (1-t)\,x_0 + t\,x_1$$

The model regresses the constant velocity of a straight path from a data sample x0 to a noise sample x1. Straight paths are why a first-order Euler step is accurate, and why the flow-matching generation reaches good images in a few dozen steps rather than hundreds.