Conditioning and classifier-free guidance
A denoiser that ignores its prompt will produce something, just not what you asked for. Conditioning attaches a text encoder to the network so the noise prediction depends on the words, and classifier-free guidance then amplifies that dependence by extrapolating away from the unconditional prediction. The amplification is a single scalar, and it controls the trade between diversity and fidelity more sharply than any other knob in the pipeline. This part shows where the conditioning enters, what the scalar does, and why pushing it too far oversaturates.
Conditioning a model
Getting words into a noise predictor
The prompt is turned into vectors by a text encoder: CLIP for its familiar embedding, T5 for longer and more compositional prompts, or a vision-language model when the model needs to reason about the caption. The encoder output is a sequence of token features, one per token, and the question is how those features reach the denoiser's spatial features.
There are two answers, and the difference matters. In cross-attention, the image features supply the queries and the text features supply keys and values, so each patch of the latent attends to the tokens it finds relevant. In joint attention, image and text tokens are concatenated into one sequence and attend to each other with shared weights; MMDiT's separate streams are a refinement of this, keeping per-modality projections but sharing the attention. Cross-attention is cheaper and was the original design; joint attention conditions more strongly and scales better with long prompts.
A row-normalised cross-attention map: six image patches against the six tokens of one prompt. Each row sums to one, so a bright cell is a patch attending to that token. The map is seeded, not trained — the shape is what matters, not the values.
Real maps are far sparser and far more interpretable than this after training: the patch containing a cat's ear lights up on the token "cat" and almost nowhere else. That sparsity is the mechanism by which a prompt steers generation, and it is also why attention maps are the usual tool for localisation and editing.
No conditioning means no guidance: with an empty prompt the model can only sample the unconditional distribution, which is exactly the reference CFG needs.
The CFG trick
Train one model to do two jobs
Classifier guidance once required a separate noisy classifier whose gradient nudged samples toward a class. Classifier-free guidance gets the same effect without the classifier by training a single model to predict both a conditional and an unconditional noise, dropping the conditioning randomly during training so the model learns the unconditional task as well. At sampling time you run the model twice on the same input — once with the prompt and once without — and combine the two predictions linearly:
$$ \varepsilon \;=\; w\,\varepsilon_\text{cond} \;+\; (1-w)\,\varepsilon_\text{uncond}. $$
When $w = 1$ this is pure conditional prediction. When $w = 0$ it is the unconditional model, a plain sample from the data distribution. Values between and beyond those two points are the interesting part: the combination is a linear extrapolation, not an interpolation, because the coefficient on the unconditional term goes negative once $w > 1$. Subtracting a multiple of the unconditional direction pushes the sample away from the generic and toward whatever makes the prompt distinctive.
In code this is Diffusion.cfg(w, uncond, cond), applied componentwise to the noise prediction at every step. It costs one extra forward pass per step, which is the entire price of guidance.
The guidance scale
Diversity collapses, then oversaturation sets in
The demo below is a two-dimensional stand-in for the whole mechanism. The data distribution is a mixture of three modes; the conditional target is a single one of them; the unconditional target is the mixture. (The unconditional predictor here is the mixture's posterior mean blended slightly toward the global mean — the small approximation error a real unconditional model carries, and the reason guidance has anything to extrapolate at all. With two perfect predictors the guided field would vanish at the mode and no scale could push past it.) Every sample starts somewhere in the mixture and is then moved by the guided prediction, so you can see in one picture what the scalar does. At $w = 0$ the samples stay spread across the mixture; near $w = 1$ they collapse onto the conditional mode; past that the extrapolated fixed point moves through and beyond the mode, and the cloud visibly overshoots it.
Top: sample cloud after guided sampling, with the three mixture modes (blue), the conditional mode (pink) and the target width (dashed). Bottom: mean distance to the conditional mode against $w$ — the U shape is diversity collapse on the left and oversaturation on the right.
The negative-prompt toggle offsets the unconditional prediction, so the guided result is pushed away from it. With $w=1$ it has no effect at all, because the unconditional term is multiplied by zero.
The readout is Diffusion.guidanceEffect, which returns the guided value alongside how far past the conditional prediction the result has been extrapolated. Watch that excess grow linearly in $w$ while the sample cloud first tightens and then scatters: that is the same oversaturation an image model shows as washed-out, over-contrasted frames at scale 15.
Rescale and intervals
Two fixes for the artifacts guidance introduces
High guidance makes the predicted noise too large, which is the mechanism behind over-saturation. Guidance rescaling measures the standard deviation of the guided prediction and scales it back toward the conditional prediction's, so the direction is preserved while the magnitude is restrained. It is the line Diffusion.cfgRescale(x, ratio), and it is why some pipelines expose a rescale parameter next to the guidance scale.
The second fix is to apply guidance only where it helps. Several models find that full guidance late in the trajectory harms fine detail, while early and mid-trajectory guidance is what fixes the composition. A guidance interval leaves the scale at a baseline outside a chosen band of timesteps and switches to the guided value inside it — the step function Diffusion.cfgInterval(wBase, w, t, tLo, tHi) plots the weight actually used.
Effective guidance weight along the trajectory for a baseline of 1 and the current scale, with the interval bounds dashed. Inside the interval the full scale applies; outside it, the baseline.
An interval of the middle 60% of the trajectory is the common setting, and it typically allows a higher guidance scale than a constant schedule would tolerate, because the distortion is confined to the steps where the composition is still being decided.
Rescale and intervals address different symptoms: rescale controls the magnitude of the extrapolation, an interval controls when it is applied.
Where this shows up
One scalar, every pipeline
Negative prompts and editing
A negative prompt is nothing more than a second unconditional run conditioned on what you do not want. Editing pipelines invert the same formula to steer toward a source image instead of a prompt, which is how diffusion-based image-to-image and inversion methods stay faithful to their input.
Guidance in time
Video models usually guide less than image models, because temporal consistency degrades quickly when the prediction is extrapolated. Audio models face the same trade with a stronger prior on silence, and commonly pair a modest scale with a negative prompt for noise.
The knob is universal because the trick is universal: run the model with and without conditioning, then extrapolate. Any conditional generative model with an unconditional fallback can be guided this way, and almost every one that ships is.
Further reading
Ho and Salimans introduced classifier-free guidance as a way to remove the auxiliary classifier; Dhariwal and Nichol then showed it beating classifier guidance outright. The rescale trick and the interval schedule come from later production experience, and the MMDiT paper is the clearest account of joint versus cross attention at scale.
Read Ho and Salimans first for the formula, Dhariwal and Nichol for the empirical case, Lin and coauthors for rescaling, and Esser and coauthors for the attention variant.
- Jonathan Ho and Tim Salimans, "Classifier-Free Diffusion Guidance", 2022 — the formula, and training one model for both conditional and unconditional prediction.
- Prafulla Dhariwal and Alex Nichol, "Diffusion Models Beat GANs on Image Synthesis", 2021 — classifier guidance and the case against it.
- Shanchuan Lin, Bingchen Liu, Jiashi Li and Xiao Yang, "Common Diffusion Noise Schedules and Sample Steps are Flawed", 2023 — variance-preservation fixes, including guidance rescaling.
- Patrick Esser and coauthors, "Scaling Rectified Flow Transformers for High-Resolution Image Synthesis", 2024 — joint attention and the MMDiT conditioning design.
Cheat sheet
| Term | Meaning here |
|---|---|
| Text encoder | CLIP, T5 or a VLM; turns the prompt into token features |
| Cross-attention | Image queries, text keys and values; how a patch reads the prompt |
| Joint attention | Image and text tokens in one sequence with shared attention weights |
| Unconditional model | The same network run with conditioning dropped; the CFG reference |
| $\varepsilon = w\varepsilon_\text{cond} + (1-w)\varepsilon_\text{uncond}$ | Classifier-free guidance as linear extrapolation |
| $w = 1$ | Pure conditional prediction; diversity preserved |
| $w \gg 1$ | Extrapolation; prompt fidelity up, diversity down, oversaturation risk |
| Rescale | Restores the guided prediction's standard deviation to control saturation |
| Guidance interval | Applies the full scale only inside a band of timesteps |
| Negative prompt | An offset applied to the unconditional prediction |