Steering a frozen model
A trained diffusion model is a fixed function of noise and conditioning, and retraining it every time you want a new subject or a new composition is out of the question. So control is bolted on around the weights instead: a tiny low-rank update, a second network that injects structure, an adapter that supplies an identity, or a noise level that decides how much of an existing image survives. This part draws where each intervention touches the trajectory, and shows why the low-noise end is where fidelity is decided.
Adapting the weights
A rank-r correction to a frozen matrix
Low-rank adaptation, LoRA, leaves the pretrained weight $W$ untouched and learns a correction $\Delta W = (\alpha/r)\,BA$, where $B$ is tall, $A$ is wide, and the product has rank at most $r$. Only $A$ and $B$ receive gradients, and the trained model is still the original one until the delta is added in. The parameter count is the selling point: a full $16\times16$ matrix has 256 entries, while a rank-4 correction of the same shape has $2\cdot4\cdot16 = 128$ — and for the large attention projections of a real model the ratio is far more lopsided, because $r$ stays small while the matrix grows.
The heatmaps below show the correction and the corrected matrix at the current rank. The left panel is the delta, which is constrained to lie in a low-rank subspace no matter how the slider moves; the right is the frozen matrix after the delta is added. At rank one the delta is a single outer product, a pattern of rows all proportional to one direction; raising the rank lets independent directions accumulate until the correction can express an arbitrary matrix.
Left: the low-rank update $\Delta W = (\alpha/r)\,BA$. Right: the frozen $W$ with the update added, on the same colour scale.
DreamBooth and textual inversion solve the same problem differently: one fine-tunes the whole model on a handful of subject images with a class-preservation loss, the other learns a new token embedding and leaves the weights alone.
Structural conditioning
ControlNet and the parallel copy
Text is a weak handle on geometry. A caption can say "a person with arms crossed" but not where the arms are, and no amount of prompt engineering fixes a pose. ControlNet answers this by training a second network that copies the encoder half of the denoiser and accepts an extra spatial input — a pose skeleton, a depth map, a set of edges, a segmentation mask — whose features are added into the frozen model through zero-initialised connections.
The zero initialisation is the trick that makes it safe: at the start of training the control branch contributes nothing, so the base model is unchanged and the new signal is learned as a residual. The conditioning bites hardest at the low-noise end, where the model is committed to layout, so a strong structural input holds the geometry while the base model fills in texture. Crucially, the base weights never move; a ControlNet for a new modality is a new branch, and several branches can be stacked for multiple constraints.
Identity and style
A vector where the structure is not spatial
Some conditioning is neither text nor geometry: it is the identity of a particular face, the look of a particular palette. IP-Adapter handles this by encoding a reference image with a small image encoder and feeding the resulting vector into the denoiser through cross-attention, exactly where text tokens are injected. Nothing in the base model changes and nothing is spatial, so it composes with a ControlNet and a prompt at the same time: structure from the skeleton, identity from the reference, semantics from the caption.
DreamBooth takes the opposite route and moves the weights, fine-tuning the whole model on a few images of a subject paired with a rare token, plus a class-preservation term that stops the model from forgetting what the class looks like. Textual inversion keeps the weights and learns only a new embedding for that token, which is cheaper but can express less. The three methods are a spectrum from a vector to a full fine-tune, and each has a characteristic failure: an adapter leaks style, a DreamBooth overfits its background, an inverted token struggles with poses it never saw.
Editing an existing image
Masks and noise levels
Editing starts from an image rather than noise, and the two classic controls are a mask and a noise level. Inpainting keeps the pixels outside the mask and lets the model regenerate inside it, so the composited result is the original everywhere the mask is off — the demo below mixes a generated gradient into a masked disc of a procedural image. The mask is a hard choice of what survives, and the model only ever sees the trajectory inside it.
SDEdit is softer. It adds noise to the whole image up to a chosen time $t_0$ and then denoises with the model, which lets the edit be large or small without a mask. The bottom canvas traces that path for several noise levels: with a little noise the result stays close to the original, and with a lot it drifts toward whatever the model considers plausible. That single knob is the difference between a retouch and a re-imagining, and it is the same dial the inversion methods tune when they search for the noise that reconstructs an image exactly.
A masked disc of the base image replaced with generated content; the dashed circle marks the mask boundary.
Denoising paths from a start at noise level $t_0$. Faint curves are other levels; the bold curve is the one selected.
Where this shows up
Every interface you have used
LoRA libraries and ControlNet packs
The community ecosystems around Stable Diffusion are exactly this part's primitives: hundreds of small rank corrections, each a style or a character, plus a ControlNet per structural modality. The latent diffusion backbone is what makes them cheap, because the adapters sit on the compressed latent rather than the pixel grid.
Edits that must stay temporally consistent
In video, an inpainting mask that moves and an IP-Adapter identity that must persist across frames are the same two ideas under a temporal constraint, developed in the video coherence chapter. Reference-image conditioning is also how subject-preserving generation in multimodal models is exposed to a user.
The unifying observation is that control is layered. A prompt sets semantics, a ControlNet sets geometry, an adapter sets identity, and a noise level sets how much of an existing image is allowed to change. Each layer touches the trajectory at a different noise range, and getting a good result is usually a matter of not asking two layers to decide the same thing.
Further reading
Control started as a way to avoid retraining and became a small field of its own. These are the five papers that define its vocabulary: the low-rank correction, the subject fine-tune, the spatial branch, the image adapter, and the noise-level edit.
Hu and coauthors show that a rank-r correction matches full fine-tuning on many tasks; Ruiz and coauthors fine-tune on a subject with a preservation loss; Zhang and coauthors add the parallel control branch; Ye and coauthors supply identity through cross-attention; and Meng and coauthors introduce the noise-level edit that needs no inversion.
- Edward Hu, Yelong Shen, Phillip Wallis and coauthors, "LoRA: Low-Rank Adaptation of Large Language Models", 2021 — the low-rank update.
- Nataniel Ruiz, Yuanzhen Li, Varun Jampani and coauthors, "DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation", 2022 — a rare token and a class-preservation loss.
- Lvmin Zhang, Anyi Rao and Maneesh Agrawala, "Adding Conditional Control to Text-to-Image Diffusion Models", 2023 — ControlNet and the zero-initialised branch.
- Hu Ye, Jun Zhang, Sibo Liu, Xiao Han and Wei Yang, "IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models", 2023 — an image encoded into a cross-attention vector.
- Chenlin Meng, Yutong He, Yang Song and coauthors, "SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations", 2021 — editing by adding noise to a chosen level.
Cheat sheet
| Method | What it changes |
|---|---|
| LoRA | Adds a rank-r update $\Delta W = (\alpha/r)BA$; base weights frozen |
| DreamBooth | Fine-tunes the whole model on a subject with a preservation loss |
| Textual inversion | Learns a new token embedding; weights untouched |
| ControlNet | A parallel encoder branch, added through zero-initialised connections |
| IP-Adapter | Encodes a reference image into a vector injected by cross-attention |
| Inpainting | Keeps pixels outside the mask, regenerates inside; Diffusion.inpaintMix |
| SDEdit | Noise to level $t_0$, then denoise; the size of the edit is $t_0$ |
| Where it bites | Structure at low noise, identity through attention, edits at chosen $t_0$ |