Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Adapting the weights

A rank-r correction to a frozen matrix

Low-rank adaptation, LoRA, leaves the pretrained weight $W$ untouched and learns a correction $\Delta W = (\alpha/r)\,BA$, where $B$ is tall, $A$ is wide, and the product has rank at most $r$. Only $A$ and $B$ receive gradients, and the trained model is still the original one until the delta is added in. The parameter count is the selling point: a full $16\times16$ matrix has 256 entries, while a rank-4 correction of the same shape has $2\cdot4\cdot16 = 128$ — and for the large attention projections of a real model the ratio is far more lopsided, because $r$ stays small while the matrix grows.

The heatmaps below show the correction and the corrected matrix at the current rank. The left panel is the delta, which is constrained to lie in a low-rank subspace no matter how the slider moves; the right is the frozen matrix after the delta is added. At rank one the delta is a single outer product, a pattern of rows all proportional to one direction; raising the rank lets independent directions accumulate until the correction can express an arbitrary matrix.

Left: the low-rank update $\Delta W = (\alpha/r)\,BA$. Right: the frozen $W$ with the update added, on the same colour scale.

DreamBooth and textual inversion solve the same problem differently: one fine-tunes the whole model on a handful of subject images with a class-preservation loss, the other learns a new token embedding and leaves the weights alone.

💡 By the end of this part you'll be able to say which knob changes the weights (LoRA, DreamBooth), which adds a parallel control network (ControlNet), which supplies a global vector (IP-Adapter), and which simply chooses a noise level on the existing trajectory (SDEdit).
2

Structural conditioning

ControlNet and the parallel copy

Text is a weak handle on geometry. A caption can say "a person with arms crossed" but not where the arms are, and no amount of prompt engineering fixes a pose. ControlNet answers this by training a second network that copies the encoder half of the denoiser and accepts an extra spatial input — a pose skeleton, a depth map, a set of edges, a segmentation mask — whose features are added into the frozen model through zero-initialised connections.

The zero initialisation is the trick that makes it safe: at the start of training the control branch contributes nothing, so the base model is unchanged and the new signal is learned as a residual. The conditioning bites hardest at the low-noise end, where the model is committed to layout, so a strong structural input holds the geometry while the base model fills in texture. Crucially, the base weights never move; a ControlNet for a new modality is a new branch, and several branches can be stacked for multiple constraints.

3

Identity and style

A vector where the structure is not spatial

Some conditioning is neither text nor geometry: it is the identity of a particular face, the look of a particular palette. IP-Adapter handles this by encoding a reference image with a small image encoder and feeding the resulting vector into the denoiser through cross-attention, exactly where text tokens are injected. Nothing in the base model changes and nothing is spatial, so it composes with a ControlNet and a prompt at the same time: structure from the skeleton, identity from the reference, semantics from the caption.

DreamBooth takes the opposite route and moves the weights, fine-tuning the whole model on a few images of a subject paired with a rare token, plus a class-preservation term that stops the model from forgetting what the class looks like. Textual inversion keeps the weights and learns only a new embedding for that token, which is cheaper but can express less. The three methods are a spectrum from a vector to a full fine-tune, and each has a characteristic failure: an adapter leaks style, a DreamBooth overfits its background, an inverted token struggles with poses it never saw.

4

Editing an existing image

Masks and noise levels

Editing starts from an image rather than noise, and the two classic controls are a mask and a noise level. Inpainting keeps the pixels outside the mask and lets the model regenerate inside it, so the composited result is the original everywhere the mask is off — the demo below mixes a generated gradient into a masked disc of a procedural image. The mask is a hard choice of what survives, and the model only ever sees the trajectory inside it.

SDEdit is softer. It adds noise to the whole image up to a chosen time $t_0$ and then denoises with the model, which lets the edit be large or small without a mask. The bottom canvas traces that path for several noise levels: with a little noise the result stays close to the original, and with a lot it drifts toward whatever the model considers plausible. That single knob is the difference between a retouch and a re-imagining, and it is the same dial the inversion methods tune when they search for the noise that reconstructs an image exactly.

A masked disc of the base image replaced with generated content; the dashed circle marks the mask boundary.

Denoising paths from a start at noise level $t_0$. Faint curves are other levels; the bold curve is the one selected.

⚠ Editing and inversion are not the same problem. SDEdit needs no inverse and no optimisation, which is why it is fast, but it cannot promise that the untouched regions stay identical. Methods based on DDIM inversion spend extra passes to recover the exact noise that reconstructs the original, buying faithfulness at low noise levels at the cost of those passes.
5

Where this shows up

Every interface you have used

Image

LoRA libraries and ControlNet packs

The community ecosystems around Stable Diffusion are exactly this part's primitives: hundreds of small rank corrections, each a style or a character, plus a ControlNet per structural modality. The latent diffusion backbone is what makes them cheap, because the adapters sit on the compressed latent rather than the pixel grid.

Video & editing

Edits that must stay temporally consistent

In video, an inpainting mask that moves and an IP-Adapter identity that must persist across frames are the same two ideas under a temporal constraint, developed in the video coherence chapter. Reference-image conditioning is also how subject-preserving generation in multimodal models is exposed to a user.

The unifying observation is that control is layered. A prompt sets semantics, a ControlNet sets geometry, an adapter sets identity, and a noise level sets how much of an existing image is allowed to change. Each layer touches the trajectory at a different noise range, and getting a good result is usually a matter of not asking two layers to decide the same thing.

Further reading

Control started as a way to avoid retraining and became a small field of its own. These are the five papers that define its vocabulary: the low-rank correction, the subject fine-tune, the spatial branch, the image adapter, and the noise-level edit.

Hu and coauthors show that a rank-r correction matches full fine-tuning on many tasks; Ruiz and coauthors fine-tune on a subject with a preservation loss; Zhang and coauthors add the parallel control branch; Ye and coauthors supply identity through cross-attention; and Meng and coauthors introduce the noise-level edit that needs no inversion.

Cheat sheet

MethodWhat it changes
LoRAAdds a rank-r update $\Delta W = (\alpha/r)BA$; base weights frozen
DreamBoothFine-tunes the whole model on a subject with a preservation loss
Textual inversionLearns a new token embedding; weights untouched
ControlNetA parallel encoder branch, added through zero-initialised connections
IP-AdapterEncodes a reference image into a vector injected by cross-attention
InpaintingKeeps pixels outside the mask, regenerates inside; Diffusion.inpaintMix
SDEditNoise to level $t_0$, then denoise; the size of the edit is $t_0$
Where it bitesStructure at low noise, identity through attention, edits at chosen $t_0$
7

Check your understanding

0/4 answered