Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Conditioning a model

Getting words into a noise predictor

The prompt is turned into vectors by a text encoder: CLIP for its familiar embedding, T5 for longer and more compositional prompts, or a vision-language model when the model needs to reason about the caption. The encoder output is a sequence of token features, one per token, and the question is how those features reach the denoiser's spatial features.

There are two answers, and the difference matters. In cross-attention, the image features supply the queries and the text features supply keys and values, so each patch of the latent attends to the tokens it finds relevant. In joint attention, image and text tokens are concatenated into one sequence and attend to each other with shared weights; MMDiT's separate streams are a refinement of this, keeping per-modality projections but sharing the attention. Cross-attention is cheaper and was the original design; joint attention conditions more strongly and scales better with long prompts.

A row-normalised cross-attention map: six image patches against the six tokens of one prompt. Each row sums to one, so a bright cell is a patch attending to that token. The map is seeded, not trained — the shape is what matters, not the values.

Real maps are far sparser and far more interpretable than this after training: the patch containing a cat's ear lights up on the token "cat" and almost nowhere else. That sparsity is the mechanism by which a prompt steers generation, and it is also why attention maps are the usual tool for localisation and editing.

No conditioning means no guidance: with an empty prompt the model can only sample the unconditional distribution, which is exactly the reference CFG needs.

💡 By the end of this part you'll be able to say what cross- and joint attention each do, write the CFG formula from memory, predict what a large guidance scale does to diversity and to saturation, and know what rescaling and guidance intervals are for.
2

The CFG trick

Train one model to do two jobs

Classifier guidance once required a separate noisy classifier whose gradient nudged samples toward a class. Classifier-free guidance gets the same effect without the classifier by training a single model to predict both a conditional and an unconditional noise, dropping the conditioning randomly during training so the model learns the unconditional task as well. At sampling time you run the model twice on the same input — once with the prompt and once without — and combine the two predictions linearly:

$$ \varepsilon \;=\; w\,\varepsilon_\text{cond} \;+\; (1-w)\,\varepsilon_\text{uncond}. $$

When $w = 1$ this is pure conditional prediction. When $w = 0$ it is the unconditional model, a plain sample from the data distribution. Values between and beyond those two points are the interesting part: the combination is a linear extrapolation, not an interpolation, because the coefficient on the unconditional term goes negative once $w > 1$. Subtracting a multiple of the unconditional direction pushes the sample away from the generic and toward whatever makes the prompt distinctive.

In code this is Diffusion.cfg(w, uncond, cond), applied componentwise to the noise prediction at every step. It costs one extra forward pass per step, which is the entire price of guidance.

⚠ Guidance is a distortion, not an improvement. The quantity being sampled is no longer the conditional distribution; it is a sharpened version of it that trades mode coverage for prompt adherence. At high $w$ the samples look more like the prompt and less like the data, which is a feature until it becomes the familiar over-saturated artifact.
3

The guidance scale

Diversity collapses, then oversaturation sets in

The demo below is a two-dimensional stand-in for the whole mechanism. The data distribution is a mixture of three modes; the conditional target is a single one of them; the unconditional target is the mixture. (The unconditional predictor here is the mixture's posterior mean blended slightly toward the global mean — the small approximation error a real unconditional model carries, and the reason guidance has anything to extrapolate at all. With two perfect predictors the guided field would vanish at the mode and no scale could push past it.) Every sample starts somewhere in the mixture and is then moved by the guided prediction, so you can see in one picture what the scalar does. At $w = 0$ the samples stay spread across the mixture; near $w = 1$ they collapse onto the conditional mode; past that the extrapolated fixed point moves through and beyond the mode, and the cloud visibly overshoots it.

Top: sample cloud after guided sampling, with the three mixture modes (blue), the conditional mode (pink) and the target width (dashed). Bottom: mean distance to the conditional mode against $w$ — the U shape is diversity collapse on the left and oversaturation on the right.

The negative-prompt toggle offsets the unconditional prediction, so the guided result is pushed away from it. With $w=1$ it has no effect at all, because the unconditional term is multiplied by zero.

The readout is Diffusion.guidanceEffect, which returns the guided value alongside how far past the conditional prediction the result has been extrapolated. Watch that excess grow linearly in $w$ while the sample cloud first tightens and then scatters: that is the same oversaturation an image model shows as washed-out, over-contrasted frames at scale 15.

⚠ Guidance and diversity trade off, always. Every unit of $w$ above one makes the samples more prompt-faithful and less varied. The standard practice is to tune it per model and per prompt length, and to accept that a high scale means fewer distinct outputs for the same seed budget.
4

Rescale and intervals

Two fixes for the artifacts guidance introduces

High guidance makes the predicted noise too large, which is the mechanism behind over-saturation. Guidance rescaling measures the standard deviation of the guided prediction and scales it back toward the conditional prediction's, so the direction is preserved while the magnitude is restrained. It is the line Diffusion.cfgRescale(x, ratio), and it is why some pipelines expose a rescale parameter next to the guidance scale.

The second fix is to apply guidance only where it helps. Several models find that full guidance late in the trajectory harms fine detail, while early and mid-trajectory guidance is what fixes the composition. A guidance interval leaves the scale at a baseline outside a chosen band of timesteps and switches to the guided value inside it — the step function Diffusion.cfgInterval(wBase, w, t, tLo, tHi) plots the weight actually used.

Effective guidance weight along the trajectory for a baseline of 1 and the current scale, with the interval bounds dashed. Inside the interval the full scale applies; outside it, the baseline.

An interval of the middle 60% of the trajectory is the common setting, and it typically allows a higher guidance scale than a constant schedule would tolerate, because the distortion is confined to the steps where the composition is still being decided.

Rescale and intervals address different symptoms: rescale controls the magnitude of the extrapolation, an interval controls when it is applied.

5

Where this shows up

One scalar, every pipeline

Image

Negative prompts and editing

A negative prompt is nothing more than a second unconditional run conditioned on what you do not want. Editing pipelines invert the same formula to steer toward a source image instead of a prompt, which is how diffusion-based image-to-image and inversion methods stay faithful to their input.

Video & audio

Guidance in time

Video models usually guide less than image models, because temporal consistency degrades quickly when the prediction is extrapolated. Audio models face the same trade with a stronger prior on silence, and commonly pair a modest scale with a negative prompt for noise.

The knob is universal because the trick is universal: run the model with and without conditioning, then extrapolate. Any conditional generative model with an unconditional fallback can be guided this way, and almost every one that ships is.

Further reading

Ho and Salimans introduced classifier-free guidance as a way to remove the auxiliary classifier; Dhariwal and Nichol then showed it beating classifier guidance outright. The rescale trick and the interval schedule come from later production experience, and the MMDiT paper is the clearest account of joint versus cross attention at scale.

Read Ho and Salimans first for the formula, Dhariwal and Nichol for the empirical case, Lin and coauthors for rescaling, and Esser and coauthors for the attention variant.

Cheat sheet

TermMeaning here
Text encoderCLIP, T5 or a VLM; turns the prompt into token features
Cross-attentionImage queries, text keys and values; how a patch reads the prompt
Joint attentionImage and text tokens in one sequence with shared attention weights
Unconditional modelThe same network run with conditioning dropped; the CFG reference
$\varepsilon = w\varepsilon_\text{cond} + (1-w)\varepsilon_\text{uncond}$Classifier-free guidance as linear extrapolation
$w = 1$Pure conditional prediction; diversity preserved
$w \gg 1$Extrapolation; prompt fidelity up, diversity down, oversaturation risk
RescaleRestores the guided prediction's standard deviation to control saturation
Guidance intervalApplies the full scale only inside a band of timesteps
Negative promptAn offset applied to the unconditional prediction
7

Check your understanding

0/4 answered