The DiT backbone
The U-Net's resolution pyramid is an assumption about images: that locality matters, that nearby pixels are more related than distant ones. A transformer refuses the assumption, cutting the latent into a flat sequence of patches and letting every token attend to every other. The Diffusion Transformer — DiT — is that idea with one addition that made it trainable: conditioning by adaptive layer norm, initialised to identity. This part patches a latent, modulates a block, and prices the whole thing against the U-Net it replaced.
Patching a latent into tokens
From a grid to a sequence
A transformer consumes a sequence, so the first job is to make one. The latent is divided into a regular grid of non-overlapping patches, each patch is flattened, and a linear layer projects it to the model's width. The result is a token sequence whose length is the number of patches, and it is on that sequence, not on pixels, that attention operates. A 32×32 latent with 4×4 patches gives 64 tokens; with 2×2 patches it gives 256.
Patch size is a resolution-versus-cost dial. Small patches keep fine spatial detail and produce a long sequence, which attention charges for quadratically; large patches are cheap and blunt. Most trained DiTs use patch size 2 on the latent, which is equivalent to patch size 16 on the 8×-compressed image.
The grid drawn on the latent is the tokenisation: each cell becomes one token before any network runs.
Plain transformer blocks
No convolutions, no pyramid, just attention and a feed-forward
Once the latent is a sequence of tokens, the denoiser becomes the standard transformer block repeated: layer norm, multi-head self-attention, a residual add, then layer norm, a feed-forward MLP, and another residual. There is no convolution anywhere and no resolution pyramid at all. Every token lives at the same scale, and every token can reach every other token in a single attention layer.
The parameter count of a block is dominated by four projection matrices of size $d\times d$ for the attention, plus two of size $d\times 4d$ for the MLP — about $12d^2$ per block, or $4d^2 + 2d\cdot 4d$ in the accounting GenMedia.cost.transformerParams uses. The compute is a different matter: attention contributes a term quadratic in the token count, while the projections and MLP are linear in it. At 256 tokens the quadratic term is small; at 4096 tokens it dominates, which is why latent compression is a prerequisite rather than a convenience.
Dropping the pyramid has a real cost. A U-Net's convolutions embed locality and translation equivariance for free; a transformer must learn them from data, which takes more parameters and more compute. In exchange the architecture is uniform, scales predictably, and can be shrunk or widened without re-deriving how the resolutions line up.
adaLN-Zero
Conditioning that starts as the identity
A transformer block needs to know the timestep and the class or prompt. The U-Net injects that information into every residual block; a DiT does the same through adaptive layer norm. A small MLP reads the conditioning vector and emits, for each block, a scale $\gamma$ and a shift $\beta$ that modulate the features after normalisation:
$$ \text{adaLN-Zero}(x) = x\,(1+\gamma) + \beta. $$
The zero in the name is the initialisation. At the start of training $\gamma$ and $\beta$ are set to zero, so $x(1+0)+0 = x$ and every conditioned block is exactly the identity. Gradients flow cleanly through a stack of identity mappings, and the modulation is learned in gradually rather than imposed on an untrained network. The same idea modulates the attention and MLP sublayers separately, and it costs only $6d$ extra parameters per block — which is why adaptive conditioning is cheap even at scale.
A fixed input vector $x$ (light bars) and the same vector after adaLN-Zero (pink bars) with $\beta = 0.4\gamma$. At $\gamma = 0$ the two coincide: the block is the identity.
Drag γ to zero and the bars overlap exactly. That overlap is the initialisation that makes deep DiTs trainable.
U-Net against DiT
Same resolution, two architectures
Put the two backbones side by side on the same 32×32 latent and the trade is visible in the parameter and FLOP counts. The U-Net's cost is spread across a pyramid and concentrates at the wide, deep levels; the DiT's cost is uniform across a stack of identical blocks and scales cleanly with depth and width. Slide the depth and width and watch the DiT bill move while the U-Net bill stays put.
Total parameters for the U-Net envelope and for a DiT with the depth and width set below, both operating on the same 32×32 latent with 2×2 patches.
The comparison is not that one architecture is smaller. It is that the DiT's cost is a smooth function of two integers, so a scaling study can sweep them independently and plot the result. The U-Net's cost is a function of many coupled choices — how many levels, how many blocks each, where attention sits — and moving one moves everything else.
Scaling and MMDiT
What uniformity buys, and where it was split again
Because a DiT is a uniform stack, its capacity is described by two numbers and its compute by one, and that is exactly the setup a scaling law needs. The DiT paper swept depth, width and patch size against measured Gflops and found that sample quality tracked compute closely, with the largest models setting the state of the art at the time. Scaling a uniform architecture is a matter of arithmetic; scaling a pyramid means re-deciding the pyramid every time.
The first large follow-ups pushed the uniform idea to its limit and then broke it deliberately. MMDiT, the backbone of Stable Diffusion 3, keeps two separate streams of tokens — one for the noisy image patches, one for the text — each with its own set of weights, and fuses them only inside the attention operation, where image tokens and text tokens are concatenated into a single sequence of keys and values. That gives the text its own representation without forcing it through the image stream, and it is why MMDiT conditions better on long prompts than a single-stream DiT with cross-attention bolted on.
Other variants trade the same way. U-ViT keeps a long skip from input to output; PixArt-α adds cross-attention to a pretrained text encoder to save training compute. The uniform block survives all of them; what changes is where the conditioning enters.
Where this shows up
The architecture the frontier settled on
SD3, Flux, Sora
MMDiT is the backbone of SD3 and Flux, and video models patchify a spatio-temporal latent into the same kind of token sequence. The guidance machinery is unchanged; only the denoiser underneath it was swapped.
The ViT inheritance
Patchify-then-attend is the Vision Transformer's move, and a DiT is a ViT that predicts noise instead of a label. Everything the transformer stack knows about scaling, mixed precision and attention kernels transfers directly.
Once a denoiser is a plain transformer, it inherits the transformer's entire toolchain: flash attention, tensor parallelism, and the ablations that say what happens when you scale depth rather than width. That inheritance, more than any single benchmark, is why the field moved.
Further reading
The DiT paper is short and unusually clean, and it is the single best source for the adaLN-Zero trick and the scaling curves. The Vision Transformer is the direct ancestor; the MMDiT paper is where the uniform stream was deliberately split.
Peebles and Xie introduce the DiT and the Gflops scaling study; Dosovitskiy and coauthors introduce the patch-embedding transformer; Esser and coauthors describe MMDiT; and Bao and coauthors give the long-skip variant.
- William Peebles and Saining Xie, "Scalable Diffusion Models with Transformers", 2022 — DiT, adaLN-Zero and the scaling law.
- Alexey Dosovitskiy and coauthors, "An Image is Worth 16x16 Words", 2020 — patch embedding and the ViT that DiT inherits.
- Patrick Esser and coauthors, "Scaling Rectified Flow Transformers for High-Resolution Image Synthesis", 2024 — MMDiT's two streams fused at attention.
- Fan Bao and coauthors, "All are Worth Words: A ViT Backbone for Diffusion Models", 2022 — U-ViT and the long-skip alternative.
Cheat sheet
| Term | Meaning here |
|---|---|
| Patchify | Split the latent into non-overlapping patches; each becomes a token |
| Token count | $(H/p)\times(W/p)$; attention cost is quadratic in it |
| DiT block | Norm, self-attention, norm, MLP, two residual adds; no convolution |
| adaLN-Zero | $x(1+\gamma)+\beta$ with $\gamma=\beta=0$ at init, so the block starts as identity |
| Cost per block | About $4d^2$ for attention projections plus $2\cdot 4d^2$ for the MLP |
| MMDiT | Separate image and text streams, fused only inside attention |
| No pyramid | One resolution, one width; locality must be learned from data |
| Scaling handle | Depth $L$, width $d$ and patch size $p$ — three integers to sweep |