Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Patching a latent into tokens

From a grid to a sequence

A transformer consumes a sequence, so the first job is to make one. The latent is divided into a regular grid of non-overlapping patches, each patch is flattened, and a linear layer projects it to the model's width. The result is a token sequence whose length is the number of patches, and it is on that sequence, not on pixels, that attention operates. A 32×32 latent with 4×4 patches gives 64 tokens; with 2×2 patches it gives 256.

Patch size is a resolution-versus-cost dial. Small patches keep fine spatial detail and produce a long sequence, which attention charges for quadratically; large patches are cheap and blunt. Most trained DiTs use patch size 2 on the latent, which is equivalent to patch size 16 on the 8×-compressed image.

32×32 latent, 4×4 patches

The grid drawn on the latent is the tokenisation: each cell becomes one token before any network runs.

💡 By the end of this part you'll be able to tokenise a latent by hand, describe a DiT block and what adaLN-Zero changes, price a DiT against a U-Net at the same resolution, and explain what MMDiT splits into two streams.
2

Plain transformer blocks

No convolutions, no pyramid, just attention and a feed-forward

Once the latent is a sequence of tokens, the denoiser becomes the standard transformer block repeated: layer norm, multi-head self-attention, a residual add, then layer norm, a feed-forward MLP, and another residual. There is no convolution anywhere and no resolution pyramid at all. Every token lives at the same scale, and every token can reach every other token in a single attention layer.

The parameter count of a block is dominated by four projection matrices of size $d\times d$ for the attention, plus two of size $d\times 4d$ for the MLP — about $12d^2$ per block, or $4d^2 + 2d\cdot 4d$ in the accounting GenMedia.cost.transformerParams uses. The compute is a different matter: attention contributes a term quadratic in the token count, while the projections and MLP are linear in it. At 256 tokens the quadratic term is small; at 4096 tokens it dominates, which is why latent compression is a prerequisite rather than a convenience.

Dropping the pyramid has a real cost. A U-Net's convolutions embed locality and translation equivariance for free; a transformer must learn them from data, which takes more parameters and more compute. In exchange the architecture is uniform, scales predictably, and can be shrunk or widened without re-deriving how the resolutions line up.

⚠ Uniform sequence, uniform cost. Because the cost is quadratic in token count and linear in width, a DiT's budget is set almost entirely by the patch size and the latent resolution chosen upstream. Getting the latent wrong cannot be fixed by adding layers.
3

adaLN-Zero

Conditioning that starts as the identity

A transformer block needs to know the timestep and the class or prompt. The U-Net injects that information into every residual block; a DiT does the same through adaptive layer norm. A small MLP reads the conditioning vector and emits, for each block, a scale $\gamma$ and a shift $\beta$ that modulate the features after normalisation:

$$ \text{adaLN-Zero}(x) = x\,(1+\gamma) + \beta. $$

The zero in the name is the initialisation. At the start of training $\gamma$ and $\beta$ are set to zero, so $x(1+0)+0 = x$ and every conditioned block is exactly the identity. Gradients flow cleanly through a stack of identity mappings, and the modulation is learned in gradually rather than imposed on an untrained network. The same idea modulates the attention and MLP sublayers separately, and it costs only $6d$ extra parameters per block — which is why adaptive conditioning is cheap even at scale.

A fixed input vector $x$ (light bars) and the same vector after adaLN-Zero (pink bars) with $\beta = 0.4\gamma$. At $\gamma = 0$ the two coincide: the block is the identity.

Drag γ to zero and the bars overlap exactly. That overlap is the initialisation that makes deep DiTs trainable.

4

U-Net against DiT

Same resolution, two architectures

Put the two backbones side by side on the same 32×32 latent and the trade is visible in the parameter and FLOP counts. The U-Net's cost is spread across a pyramid and concentrates at the wide, deep levels; the DiT's cost is uniform across a stack of identical blocks and scales cleanly with depth and width. Slide the depth and width and watch the DiT bill move while the U-Net bill stays put.

Total parameters for the U-Net envelope and for a DiT with the depth and width set below, both operating on the same 32×32 latent with 2×2 patches.

The comparison is not that one architecture is smaller. It is that the DiT's cost is a smooth function of two integers, so a scaling study can sweep them independently and plot the result. The U-Net's cost is a function of many coupled choices — how many levels, how many blocks each, where attention sits — and moving one moves everything else.

5

Scaling and MMDiT

What uniformity buys, and where it was split again

Because a DiT is a uniform stack, its capacity is described by two numbers and its compute by one, and that is exactly the setup a scaling law needs. The DiT paper swept depth, width and patch size against measured Gflops and found that sample quality tracked compute closely, with the largest models setting the state of the art at the time. Scaling a uniform architecture is a matter of arithmetic; scaling a pyramid means re-deciding the pyramid every time.

The first large follow-ups pushed the uniform idea to its limit and then broke it deliberately. MMDiT, the backbone of Stable Diffusion 3, keeps two separate streams of tokens — one for the noisy image patches, one for the text — each with its own set of weights, and fuses them only inside the attention operation, where image tokens and text tokens are concatenated into a single sequence of keys and values. That gives the text its own representation without forcing it through the image stream, and it is why MMDiT conditions better on long prompts than a single-stream DiT with cross-attention bolted on.

Other variants trade the same way. U-ViT keeps a long skip from input to output; PixArt-α adds cross-attention to a pretrained text encoder to save training compute. The uniform block survives all of them; what changes is where the conditioning enters.

6

Where this shows up

The architecture the frontier settled on

Image & video

SD3, Flux, Sora

MMDiT is the backbone of SD3 and Flux, and video models patchify a spatio-temporal latent into the same kind of token sequence. The guidance machinery is unchanged; only the denoiser underneath it was swapped.

Sequence models

The ViT inheritance

Patchify-then-attend is the Vision Transformer's move, and a DiT is a ViT that predicts noise instead of a label. Everything the transformer stack knows about scaling, mixed precision and attention kernels transfers directly.

Once a denoiser is a plain transformer, it inherits the transformer's entire toolchain: flash attention, tensor parallelism, and the ablations that say what happens when you scale depth rather than width. That inheritance, more than any single benchmark, is why the field moved.

Further reading

The DiT paper is short and unusually clean, and it is the single best source for the adaLN-Zero trick and the scaling curves. The Vision Transformer is the direct ancestor; the MMDiT paper is where the uniform stream was deliberately split.

Peebles and Xie introduce the DiT and the Gflops scaling study; Dosovitskiy and coauthors introduce the patch-embedding transformer; Esser and coauthors describe MMDiT; and Bao and coauthors give the long-skip variant.

Cheat sheet

TermMeaning here
PatchifySplit the latent into non-overlapping patches; each becomes a token
Token count$(H/p)\times(W/p)$; attention cost is quadratic in it
DiT blockNorm, self-attention, norm, MLP, two residual adds; no convolution
adaLN-Zero$x(1+\gamma)+\beta$ with $\gamma=\beta=0$ at init, so the block starts as identity
Cost per blockAbout $4d^2$ for attention projections plus $2\cdot 4d^2$ for the MLP
MMDiTSeparate image and text streams, fused only inside attention
No pyramidOne resolution, one width; locality must be learned from data
Scaling handleDepth $L$, width $d$ and patch size $p$ — three integers to sweep
8

Check your understanding

0/4 answered