Generative Media, Interactively
Volume I of the Multimodal & Generative Media arc - from a distribution you never see, through the forward process and the learned denoiser, to flow matching, latent diffusion, DiT backbones, guidance and control, then video, audio, speech and the honesty problem of evaluating it all.
Every image this guide makes starts as noise. That is not a metaphor — it is the training objective. A model is shown a picture with a little noise added, over and over, until it can name the noise in any picture, and then the naming is run backwards from a field of static until something recognisable appears. This volume builds that idea from a dataset you draw with a mouse, through the denoiser you train in the browser, up to the flow-matching transformers that produced the 2026 state of the art, and then out past images into video and sound.
Nothing here assumes prior generative-model knowledge, but it does assume the foundation volumes: the probability guides for what a distribution is and what a conditional Gaussian looks like, and Probability in Action, Part 20 for the one-page diffusion on-ramp and the VAE this guide re-derives rather than assumes. The linear algebra guide supplies the low-rank decomposition behind the latent space and the calculus volume supplies the gradient flow that makes the denoiser trainable. Its sibling volume is Multimodal Models, Interactively, which starts where this one leaves off: models that read images and speech as well as make them. Where this guide needs a serving number — a latency, a shard, a cost — it points at LLM Serving, Interactively rather than repeating it, and its 3-D cousin NeRF and Gaussian splatting live in multi-view geometry.
The parts
Sampling from a distribution you never see: a 2D dataset you draw yourself, and the line between memorising it and modelling it.
The forward process on a procedurally-generated image, q(x_t|x_0) in closed form, and what ᾱ, SNR and the schedule choice actually do.
ε-, x₀- and v-prediction, and a real tiny denoiser trained in the browser with its loss curve live.
The reverse loop from noise to sample; ancestral DDPM against deterministic DDIM, and a step-count slider that shows the trade.
The score field as arrows over a density, Langevin dynamics, the SMLD to DDPM equivalence and the VE versus VP SDEs.
Curved paths against straight ones, the velocity-field view of an ODE, and why the frontier trainers made flow matching the default.
The VAE compressor that shrinks an image before it is diffused, the cost of an eight-times downsample, and the pixel versus latent budget.
The resolution pyramid and its skips, residual blocks, where self- and cross-attention sit, and interactive parameter counts.
Patching a latent into tokens, plain transformer blocks with adaLN-Zero conditioning, and U-Net against DiT.
Text encoders and cross- against joint attention, CFG as extrapolation, and what guidance scale costs in diversity.
Euler, Heun and multistep solvers against the ancestral sampler, timestep spacing as a design choice, and step budget against error.
Progressive, consistency, adversarial and shortcut distillation, and what collapsing fifty steps into one or four costs in diversity.
LoRA and DreamBooth, ControlNet and IP-Adapter, inpainting and SDEdit editing, and where along the trajectory each intervention bites.
The causal 3-D VAE, spatiotemporal patches, 3-D against factorised attention, and the token budget a clip really costs.
Image and keyframe conditioning, temporal cascades, autoregressive rollout against full-sequence diffusion, drift, and the world-model claim.
A waveform, its sample rate, the STFT and the mel scale, and the token-rate arithmetic that makes raw waveform generation hopeless.
The encoder, residual vector quantiser and decoder that turn sound into a few hundred discrete tokens a second.
From alignment and duration models to codec language models and two-stage synthesis, with cloning and streaming first-chunk latency.
CTC, RNN-T and Whisper, why streaming is hard and the two-pass fix, then voice activity, barge-in and the realtime duplex budget.
FID, FVD and CLIPScore and what each misses, preference models and Elo arenas, memorisation, and provenance under the EU AI Act.