Generative Media, Interactively — Reference: Glossary & numbers to know
This is the generative-media guide's back matter. The first table is an alphabetical index that links every term to the part that introduces it, and it is filterable: type a fragment and the rows that do not match disappear. After it come the three tables the series cites — the model lineage from latent U-Nets to flow-matching transformers, the sampler comparison that decides how many steps a picture costs, and the formulas behind the schedules, the DDIM step, guidance and the flow-matching loss. The serving-side counterpart, with the latency, memory and capacity arithmetic, lives in the LLM Serving glossary.
Glossary
Every term, linked to the part that introduces it
| Term | What it means | Introduced in |
|---|---|---|
| A | ||
| AdaLN-Zero | Conditioning that lets a DiT block stay a plain transformer: the timestep and text embeddings produce scale, shift and gate values, and the modulating projections are initialised to zero so each block starts as the identity. | Part 9 |
| Aliasing | A frequency above the Nyquist limit folding back and appearing as a lower one; the artefact that makes a low-pass filter mandatory before any downsampling. | Part 16 |
| Ancestral sampler | A sampler that injects fresh noise at every step, so the trajectory is stochastic; the original DDPM reverse process is its canonical form. | Part 4 |
| Attention (cross against joint) | Two ways to inject conditioning: a separate cross-attention pass against the prompt embedding, or text and image tokens concatenated into one sequence that self-attends jointly. | Part 10 |
| B | ||
| Backbone | The network that predicts the noise, the clean image or the velocity: a U-Net with a resolution pyramid, or a transformer over patched latents. | Part 8 |
| Barge-in | A user interrupting a spoken reply. A duplex system has to detect it and stop speaking inside the interaction budget, which is a latency problem rather than an accuracy one. | Part 19 |
| Bitrate | Bits per second a codec emits: frame rate times codebooks times bits per codebook, the number that makes a codec a token budget rather than an audio format. | Part 17 |
| C | ||
| C2PA | Content Credentials: a cryptographically signed manifest attached to a file that records who made it, with which tool and what edits. Strong while it is attached, and gone the moment a platform strips metadata. | Part 20 |
| Causal 3-D VAE | A video compressor whose temporal convolutions see only the current and past frames, so no latent frame can depend on the future. | Part 14 |
| Classifier-free guidance (CFG) | Training with the condition dropped at random, then extrapolating between the conditional and unconditional predictions at sampling time; the guidance scale is the extrapolation factor. | Part 10 |
| CLIP | The contrastive text-image encoder used both as the text encoder in the early stack and as the scoring model behind CLIPScore. | Part 10 |
| CLIPScore | The cosine between a CLIP image embedding and a text embedding, used as a reference-free alignment metric. It measures prompt adherence and is blind to realism, anatomy and aesthetics. | Part 20 |
| Codebook | The learned vector table a quantiser snaps its input to. Residual quantisation stacks several codebooks so each stage corrects what the previous one left behind. | Part 17 |
| Consistency distillation | Training the model to map any point on a sampling trajectory straight to its endpoint, which is what makes one- and few-step generation possible. | Part 12 |
| ControlNet | A trainable copy of the backbone driven by a conditioning image such as edges, depth or pose, injected back into the frozen model's blocks. | Part 13 |
| CTC | Connectionist temporal classification: an alignment-free output layer that permits a blank symbol and collapses repeats, producing monotonic text from a frame sequence. | Part 19 |
| D | ||
| DDIM | A deterministic, non-ancestral sampler built from the same trained model; it removes the per-step noise injection and allows larger jumps between timesteps. | Part 4 |
| DDPM | The denoising diffusion probabilistic model: a fixed forward noising chain plus a learned reverse chain trained on the variational bound. | Part 4 |
| Diffusion | The family of generative models that destroy data with noise and learn to reverse the destruction one step at a time. | Part 1 |
| DiT | A diffusion transformer: the latent is patched into tokens and processed by plain transformer blocks conditioned through adaptive layer norm. | Part 9 |
| Distillation (diffusion) | Training a few-step student from a many-step teacher, by progressive halving, consistency objectives, adversarial losses or shortcut models. | Part 12 |
| DPM-Solver++ | A multistep exponential-integrator solver for the probability-flow ODE; the usual default when a strong epsilon- or x0-prediction model has to run in a few dozen steps. | Part 11 |
| Duplex | Full-duplex speech: listening while speaking, which turns turn-taking and interruption into latency budgets rather than transcription problems. | Part 19 |
| E | ||
| Elo | A rating scale fitted to pairwise votes, usually through a Bradley-Terry model. Relative, prompt-set dependent, and moved by voter style preference. | Part 20 |
| Epsilon-prediction | Training the denoiser to output the noise that was added rather than the clean image; the original diffusion parameterisation and still the most common. | Part 3 |
| Euler solver | The simplest ODE integrator, one function evaluation per step; the k-diffusion and flow-matching default because straight paths make a first-order method accurate. | Part 11 |
| F | ||
| FID | Frechet Inception Distance: the Frechet distance between Gaussians fitted to real and generated feature clouds. It measures marginal realism, cannot see the conditioning, and is biased upward at small sample counts. | Part 20 |
| Flow matching | Training a velocity field that carries a simple distribution to the data along a chosen path, usually a straight line between noise and sample. | Part 6 |
| Frame rate (codec) | Codec frames per second, around fifty to seventy-five in modern audio codecs; one of the two factors of the token rate an audio language model sees. | Part 17 |
| FVD | Frechet Video Distance: FID computed on video features, so spatial and temporal quality collapse into a single scalar with wide error bars. | Part 20 |
| G | ||
| Gaussian posterior | The closed-form conditional that makes the forward process reversible in one step: given a noisy image and the clean one, the previous state is Gaussian. | Part 2 |
| GenAI-Arena | A human pairwise-vote arena for generative models whose leaderboard is an Elo fit over the votes that users cast on submitted prompts. | Part 20 |
| Guidance scale | The weight on the conditional minus unconditional difference; raising it improves prompt adherence and costs diversity and can oversaturate the image. | Part 10 |
| H | ||
| Heun solver | A second-order predictor-corrector ODE solver: one extra function evaluation per step buys substantially lower discretisation error. | Part 11 |
| Human Preference Score (HPS) | A preference model trained on the HPD v2 dataset of human choices, used as a differentiable proxy for taste. It inherits the annotator pool and is gamed once optimised against. | Part 20 |
| I | ||
| Inpainting | Regenerating a masked region while keeping the rest of the image, done by mixing noise into the latent inside the mask at every step and keeping the known region consistent. | Part 13 |
| IP-Adapter | A small adapter that injects a reference image as an extra conditioning signal into a frozen backbone, so a photograph can steer style or subject. | Part 13 |
| K | ||
| Keyframe conditioning | Supplying first, last or intermediate frames to a video model so a clip starts and ends where the user wants it to. | Part 15 |
| L | ||
| Langevin dynamics | Sampling by following the score field with a noisy gradient step; the ancestral sampler is one discretisation of it. | Part 5 |
| Latent | The compressed representation a VAE produces, about eight times smaller along each spatial axis, in which diffusion actually runs. | Part 7 |
| LoRA | A low-rank update to frozen weights; the standard way to personalise or steer a model without fine-tuning all of it. | Part 13 |
| M | ||
| Mel scale | A perceptual frequency scale that spaces bins by how they sound rather than by hertz; the vertical axis of the spectrogram most audio models consume. | Part 16 |
| Memorisation | A model reproducing a training example rather than generating a new one, concentrated in duplicated and atypical images and measured by nearest-neighbour similarity. | Part 20 |
| N | ||
| Noise schedule | The sequence of betas that sets how fast signal is destroyed, summarised by its cumulative product alpha-bar. | Part 2 |
| Nyquist frequency | Half the sample rate: the highest frequency a given rate can represent without aliasing. | Part 16 |
| O | ||
| ODE sampler | A sampler that integrates the probability-flow ODE deterministically, in contrast to the stochastic ancestral sampler; the family Euler, Heun and DPM-Solver++ belong to. | Part 11 |
| One-step generation | Producing an image in a single forward pass, by distillation or by training a new model; the extreme end of the steps-against-quality trade. | Part 12 |
| P | ||
| Patchify | Cutting a latent into fixed tiles and flattening them into a token sequence; the operation that turns an image into transformer input. | Part 9 |
| PickScore | A preference model trained on the Pick-a-Pic dataset of user choices. It correlates with human judgement better than CLIPScore and is gamed once it becomes a training signal. | Part 20 |
| Progressive distillation | Halving a sampler's step count repeatedly, each student learning to reproduce two teacher steps in one. | Part 12 |
| Provenance | Evidence about how a piece of media was made: a watermark carried in the pixels, a signed manifest in the metadata, or both. Absence of either is not evidence of authenticity. | Part 20 |
| Q | ||
| Quantization (serving) | Storing weights and activations at FP8, INT8 or FP4 to cut bytes and raise the arithmetic rate. The accuracy bill has to be measured on images rather than inherited from language-model results. | Part 20 |
| R | ||
| Rectified flow | Flow matching on straight paths plus a reflow step that straightens them further, so a first-order solver is accurate at few steps. | Part 6 |
| Residual vector quantisation (RVQ) | Quantising the residual of the previous stage across several codebooks, so reconstruction error falls with each added stage. | Part 17 |
| RNN-T | A transducer that combines a prediction network, an encoder and a joint network to emit tokens while streaming, without CTC's frame-independence assumption. | Part 19 |
| S | ||
| Sampler | The procedure that walks the trained model from noise to a sample: ancestral, a first-order ODE step, or a higher-order multistep solver. | Part 11 |
| SDEdit | Editing by adding noise to an image up to a chosen level and denoising back, which sets how far the edit is allowed to travel from the original. | Part 13 |
| Score function | The gradient of the log density with respect to the data; the quantity the denoiser implicitly estimates and the object Langevin sampling follows. | Part 5 |
| Signal-to-noise ratio (SNR) | The ratio of signal variance to noise variance at a timestep; its logarithm is the natural axis for schedules and for solver error. | Part 2 |
| Spatiotemporal patch | A latent patch spanning several frames; the token unit of a video transformer after the causal 3-D VAE has compressed space and time. | Part 14 |
| Step caching | Reusing backbone features across adjacent sampler steps and skipping the recomputation where the change is small; the cheapest server-side speedup. | Part 20 |
| SynthID | A sampling-time watermark whose key-dependent residual is spread over the whole image and recovered by a detector that holds the key. | Part 20 |
| T | ||
| Temporal cascade | A base model at low resolution plus a temporal upsampler, so clip length grows without making attention quadratic in the whole clip. | Part 15 |
| Text encoder | The frozen language or CLIP model that turns the prompt into the conditioning sequence the backbone attends to, and the reason prompt length costs prefill time. | Part 10 |
| Text to speech (TTS) | Turning text into audio: alignment or duration prediction first, then acoustic tokens or a spectrogram, then a vocoder. | Part 18 |
| Token rate | Audio tokens per second: frame rate times codebooks. It is the number that decides whether a language model over audio is tractable at all. | Part 17 |
| U | ||
| U-Net | A convolutional backbone with a resolution pyramid and skip connections; the image-generation workhorse before transformers took over the latent. | Part 8 |
| V | ||
| VAE | The variational autoencoder whose encoder compresses an image into a latent and whose decoder turns a latent back into pixels. | Part 7 |
| v-prediction | Predicting a velocity-like combination of the clean image and the noise, which behaves better than epsilon-prediction at the extremes of the schedule. | Part 3 |
| Velocity field | The field a flow-matching model learns: the direction and speed of the path that carries noise to data. | Part 6 |
| Voice activity detection (VAD) | Deciding when a person is speaking; the trigger that ends a turn and starts the recogniser in a duplex system. | Part 19 |
| Voice cloning | Reproducing an unseen speaker from a short reference recording: zero-shot with a codec language model, or few-shot with an adapter. | Part 18 |
| Vocoder | The component that turns a spectrogram or codec tokens back into a waveform, in current systems often a small diffusion or flow model of its own. | Part 18 |
| W | ||
| Watermark | A signal embedded in the output so it can be identified later, either distributed through the pixels or declared in a signed manifest. | Part 20 |
| World model | The claim that a video generator has learned dynamics that can be controlled and queried rather than merely rendered. | Part 15 |
| X | ||
| x0-prediction | Parameterising the denoiser to output the clean image directly: convenient for clipping, but badly scaled at high noise. | Part 3 |
| Z | ||
| Zero-shot cloning | Cloning a voice from a reference clip with no per-speaker fine-tuning, which is what codec language models made routine. | Part 18 |
Model lineage
From a 512-pixel latent U-Net to a flow-matching transformer
Read the table as a change of objective and of backbone rather than as a list of releases. The first generation diffused a latent with a U-Net and predicted noise; the current generation predicts a velocity along a straight path with a transformer, and the only reason the two can be discussed together is that both are trained by predicting something about a noised input and both are sampled by walking backwards. Where a parameter count or a training resolution is undisclosed, the cell says so rather than guessing.
| Model | Released | Resolution | Parameters | Objective | Backbone |
|---|---|---|---|---|---|
| Stable Diffusion 1.5 | 2022-10 | 512 px square | 0.86 B U-Net, 0.08 B VAE, CLIP ViT-L text encoder | DDPM epsilon-prediction, classifier-free guidance | Latent U-Net |
| Stable Diffusion XL | 2023-07 | 1024 px square | 2.6 B U-Net, two text encoders (CLIP ViT-L and OpenCLIP ViT-bigG) | Epsilon-prediction with size and crop conditioning; base plus refiner | Latent U-Net |
| Stable Diffusion 3.5 | 2024-10 | 1024 px, up to 1440 px | 8 B Large, 2.5 B Medium | Rectified flow with logit-normal timestep sampling | MMDiT |
| FLUX.1 | 2024-08 | 1024 to 2048 px | 12 B | Rectified flow matching; a guidance-distilled schnell variant runs in a few steps | MMDiT, double and single stream blocks |
| Qwen-Image | 2025-08 | 1024 px and above, arbitrary aspect ratio | 20 B | Flow matching, with strong text rendering | MMDiT |
| Veo 3 | 2025-05 | 720p to 1080p video at 24 fps, native audio | Undisclosed | Flow matching over video and audio latents | DiT with factorised attention |
| Sora 2 | 2025-09 | Up to 1080p video with synchronised audio | Undisclosed | Spatiotemporal latent diffusion with flow matching | DiT over spatiotemporal patches |
Released dates are public announcement dates; parameter counts are the published configuration at release, and undisclosed means the developer has not published a figure. Text encoders and VAEs are counted separately from the diffusion backbone where the counts were published separately.
Sampler comparison
Deterministic or not, what order, and what it costs
The sampler is where the step budget is decided, and the whole table follows from two questions: does the update inject noise, and how many function evaluations does one step use. Ancestral sampling injects noise, so it is stochastic and needs many small steps. The ODE solvers do not, so they can take larger jumps, and a higher-order method spends more model calls per step to make those jumps accurate. Flow-matching models make the straightest paths, which is why a first-order Euler step is competitive for them at step counts that a curved DDPM path could never reach.
| Sampler | Deterministic? | Order | Typical steps | Notes |
|---|---|---|---|---|
| DDPM ancestral | No | First | 250 to 1000 | Adds a fresh noise draw at every step, matching the training objective. The reference behaviour, and by far the most model calls. |
| DDIM | Yes when eta is zero | First | 20 to 100 | The same trained model with the per-step noise removed; supports skipping timesteps, and the eta knob interpolates back toward ancestral sampling. |
| DPM-Solver++ | Yes | Second to third | 15 to 30 | A multistep exponential integrator in log-SNR. The usual default for epsilon- and x0-prediction models, and the point where a second order pays for itself. |
| Heun (predictor-corrector) | Yes | Second | 25 to 50 | One extra function evaluation per step to correct the Euler prediction. Used by the EDM formulation and cheap at low model scales. |
| Euler | Yes | First | 20 to 50 on a curved path, far fewer on a straight one | The baseline ODE integrator: one model call per step, no memory of previous steps. Its accuracy depends entirely on how straight the path is. |
| Flow matching, Euler ODE | Yes | First | 20 to 50, or 1 to 4 after distillation | The current frontier default: straight paths make a first-order step accurate, and a guidance-distilled or consistency student can collapse the count to a handful. |
Step counts are the ranges reported for the samplers in practice, not hard limits: quality depends on the model, the schedule spacing and the guidance scale. A distilled model changes the row's typical count rather than its order.
Formulas to keep
Four expressions the guide builds everything from
The schedule fixes how much signal is left at each timestep, the DDIM step is the deterministic update written in terms of that schedule, guidance is an extrapolation between two predictions of the same model, and the flow-matching loss is what the frontier trainers replaced the noise-prediction loss with. Everything else in the series is a consequence of these.
The noise schedule
Beta is the per-step variance of the forward process and alpha-bar is its cumulative signal retention. The second expression is the one-step forward kernel: a sample at any timestep is a closed-form mixture of the clean image and a standard normal draw, which is what makes training a single-step denoiser possible.
The DDIM step
Estimate the clean image from the model's noise prediction, then re-noise it to the next timestep's level. No noise is added, so the trajectory is deterministic and the same seed and prompt give the same image twice.
Classifier-free guidance
One model is evaluated twice, with and without the condition, and the difference is scaled by the guidance weight w. At w equal to one it is ordinary conditional sampling; above one prompt adherence improves while diversity falls and colours oversaturate.
Flow-matching loss
The model regresses the constant velocity of a straight path from a data sample x0 to a noise sample x1. Straight paths are why a first-order Euler step is accurate, and why the flow-matching generation reaches good images in a few dozen steps rather than hundreds.