Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The interface between towers

Two frozen spaces and one learned adapter

A vision-language model starts from two pretrained parts that were trained separately. Pretrained means someone already spent the compute and the data to train each part on its own task — one to recognise images, one to predict text — and we get to reuse the result instead of starting over. The vision tower is a patch transformer — the encoder of the previous parts — and it emits one vector per patch at a width of, say, $1024$. Width here means how many numbers are packed into each vector. The language model expects vectors of width $4096$, and it will only accept them as the same input channel it uses for text. The bridge is whatever maps one to the other, and it has to do so for a whole grid of patch vectors at once, because the vision tower hands over every patch rather than one summary of the image. Here is the important structural choice, and it is worth pausing on because it explains almost everything else in this part: the two towers are usually frozen, or nearly so, during training, and only the bridge between them is updated. Frozen means a tower's weights are held fixed — they are read but never changed — so no training signal flows back into it. Why does that matter? Training a vision encoder from scratch, or a language model from scratch, takes enormous compute and enormous data — that is exactly why pretrained versions of each already exist and are reused here. If you unfroze both towers and trained everything jointly, you would be re-paying that entire cost just to teach the pair to talk to each other. Freezing them instead means you only ever have to train the (comparatively tiny) bridge — a translator that has to learn to speak two already-fluent languages to each other, not learn either language from scratch. That is the single fact that makes vision-language training affordable at all.

Four designs account for nearly everything in the literature. A linear projector is a single learned matrix, cheap and surprisingly strong; it is called linear because all it can do is a weighted sum, with no bend or curve in the mapping. An MLP projector adds a hidden layer and a nonlinearity, which is what LLaVA settled on. A hidden layer is an extra stage of computation between the input vector and the output vector, and a nonlinearity is a function that lets that stage bend its input rather than only rescale it — without one, stacking layers would collapse back into a single matrix. A Q-Former, from BLIP-2, uses a small transformer with a fixed set of learned query tokens that read from the image, which compresses hundreds of patches into a handful of vectors; a query token here is a learned vector that asks a question of the image features, much as a search term asks a question of a database. Gated cross-attention, from Flamingo, inserts attention layers into the language model itself so the text can look at the image without any image tokens entering the sequence at all. Cross-attention is attention in which the queries come from one stream, here the text, and the keys and values come from another, here the image, so one stream can read the other without the two being merged into a single sequence. The gate is a learned multiplier that starts at zero, which means each inserted layer is switched off at the beginning and only turns on as training finds it useful.

Everything else in this part hangs off that menu. The demo below prices the bridges: pick a design and a vision tower, and watch the trainable parameter count move. Parameter count is a rough proxy for two things at once, how much memory the bridge needs and how much data it will take to train, which is why it is the number people usually quote first. The gap between the cheapest and the most expensive option is two orders of magnitude, and that gap is the reason the choices are made the way they are — a team with a modest compute budget and a team with a large one will not reach the same conclusion from the same menu, and both can be right for their setting.

💡 By the end of this part you'll be able to name the four bridge designs, estimate what each trains, say why a Q-Former compresses a variable number of patches into a fixed number of tokens, explain how gated cross-attention keeps the language sequence free of image tokens, and say where each option puts the image in the sequence. Underneath all five is one recurring question: which parts move during training, and which stay frozen.
2

Projectors, linear and otherwise

The smallest thing that could possibly work

The linear projector is a single matrix $W$ that takes a vision vector of width $d_v$ to a language vector of width $d_l$. It has $d_v \cdot d_l$ parameters, it is trained on image-caption or instruction data, and it is applied to every patch token independently. That is the whole idea. What makes it work is that the vision tower has already been aligned to language by a contrastive objective, so the geometry is roughly right and the matrix only has to rotate and rescale it. In other words, the contrastive stage already pushed image vectors and word vectors into comparable positions, so the projector's job is closer to translation than to re-invention: it moves the vectors to the right place without having to decide afresh what they mean. Adding a hidden layer — an MLP with a nonlinearity, the LLaVA design — costs another $d_l^2$ parameters and buys a slightly more expressive map. A hidden layer is an intermediate vector the network computes before the output, and a nonlinearity is a bending or squashing function applied to it; together they let the projector fit a curved mapping rather than only a straight one, which is the sense in which it is more expressive. The MLP is worth its cost at $d_l = 4096$ and is not worth much more than that; three-layer projectors have been tried and rarely win.

Each of these options is applied per token, so the number of trainable parameters does not depend on how many patches the image produces. That is the property to hold on to: a projector's parameter count is set by the two widths, while its token count is set by the patch grid. The two budgets are separate levers. A taller image or a finer patch size will hand the language model more rows to read, but it will not add a single trainable weight to the projector, because the same small matrix is reused on every row. This is why the rest of the part can discuss a parameter budget and a token budget almost independently: one is fixed by the widths you chose, the other by the resolution you feed in.

Trainable parameters for each bridge at the selected vision tower, drawn with Guide.drawBars. The highlighted bar is the current choice; the language model itself is assumed frozen.

Estimates for a 4096-wide language model and a 768-wide Q-Former; the cross-attention bar counts four inserted gated layers. Treat them as orders of magnitude, not exact counts.

⚠ Parameter count is not the only cost. The projector's parameters are tiny next to the towers, but its activations are not: one image becomes hundreds or thousands of tokens, and every one of them is a row of the language model's attention. Activations are the intermediate vectors a network computes as it runs, held in memory so gradients can be computed; they scale with the number of tokens, not with the number of weights. The next part is about that token budget, which usually dominates the parameter budget in practice.
3

Querying the image with a Q-Former

A fixed number of queries, however big the image

BLIP-2's Q-Former is a small transformer with a fixed set of learned query tokens — typically $32$ of them — that cross-attend to the vision tower's patch features. Cross-attention means the queries come from one place and the keys and values from another: here the queries are the learned tokens and the keys and values are the patches. The queries do not depend on the image; they are parameters. What depends on the image is the attention, so each query ends up pulling a different weighted average of the patches, and the $32$ output vectors become the $32$ image tokens the language model sees. Picture a panel of $32$ specialists, each with a fixed interest, all reading the same page of patches: one tracks colour, another tracks shapes on the left, and each reports the average of whatever caught its attention. The trick is that the number $32$ is chosen by the designer, not by the resolution: a Q-Former turns an image of any size into the same short summary, which keeps the language model's context budget predictable and makes the whole system resolution-agnostic at the interface. The context budget is the number of tokens the language model can hold at once; because it is fixed, a bridge that always emits the same number of image tokens fits a fixed budget no matter how large the picture is.

The cost is that a summary is lossy. Thirty-two vectors have to carry everything the language model will ever know about the picture, and fine spatial detail — the serial number on a label, the position of a small object — is exactly what a soft average of patches destroys. The reason is that averaging is a smoothing operation: if one patch holds a tiny, unique mark and its neighbours hold background, the weighted average keeps the background and washes the mark out. The heatmap below shows the query-to-patch attention: each row is a query, each column a patch, and a sharp row is a query that has found something specific to report. The temperature slider controls how sharply each query commits; a low temperature concentrates a row onto a few patches, and a high temperature spreads it across the whole image.

Query-to-patch attention, softmaxed per row. Rows are the 32 learned queries, columns are image patches; the matrix comes from a seeded stream, so it is the same on every load.

A fixed query set means a fixed token count. That is the point of the design and also its ceiling.

💡 Resampling is compression with a knob. The Q-Former's query count is a dial between "cheap, short context" and "expressive, long context", and it is the ancestor of the perceiver-resampler used by later models; turning the dial up keeps more detail and costs more context, and turning it down does the reverse. When you see a multimodal model advertise a fixed number of visual tokens regardless of image size, this is usually the mechanism.
4

Cross-attention instead

Leave the language sequence alone

Flamingo took a different route. Rather than converting patches into tokens that sit in the language model's input, it inserted new cross-attention layers into the frozen language model's blocks, so that the text tokens can attend to the image features at several depths. There are no image tokens in the sequence at all: the text is generated exactly as before, except that every so often it is allowed to look at a set of image features that never enter the self-attention. Each inserted layer carries a learnable gate initialised at zero, so training starts from the frozen model's behaviour and gradually opens the tap — a stabiliser that matters when the rest of the network is frozen and the new layers see gradients for the first time. Gradients are the signals training uses to nudge each weight; a fresh, randomly initialised layer produces large and noisy ones, and if those signals could flow back into the pretrained text model they would disturb abilities that already work. Starting the gate at zero means the new layer contributes nothing at first, so the model's outputs are unchanged and the gradients stay small until the gate has learned something worth adding.

The advantage is that the language model's sequence length is untouched, so nothing in its positional or context story has to change, and the frozen text ability is protected. That matters because a language model has a fixed maximum sequence length and a fixed sense of where each position sits; adding hundreds of image tokens pushes against both, and a bridge that inserts nothing avoids the problem entirely. The disadvantage is that you now have trainable layers inside the language model rather than beside it, which is a larger and more intrusive edit: you are modifying the model's own blocks, so the bridge has to be built for that specific architecture rather than attached from outside. The heatmap below is text-token attention over image tokens with the gate applied; drag the gate down and the image contribution fades toward the frozen text-only model.

Effective cross-attention of text tokens (rows) over image tokens (columns), scaled by the gate. At gate zero the inserted layers are transparent and the model falls back to its text-only behaviour.

The gate starts at zero and is learned. It is the standard answer to "how do I add a new pathway to a frozen model without wrecking it".

5

Where the image tokens go

One prefix, one bridge, no free lunch

Once a bridge is chosen, the resulting image representation has to be spliced into the language model's input, and where it lands is a design decision with consequences. LLaVA concatenates the projected image tokens into the sequence before the instruction text, so the language model sees a row of vectors that stand for the picture followed by the words. That chunk of image vectors sitting at the front of the input is called a prefix, and the model has no way to tell it apart from ordinary text except by where it sits and what it contains; the words come after it and attend to it like any other token. The image tokens then attend to the text and the text to the image, both through the ordinary self-attention of the frozen language model, and the number of extra positions is exactly the number of image tokens. A Q-Former inserts its fixed query set at the same place, in the same way, just far fewer of them. Gated cross-attention inserts nothing: the image is a side channel, delivered to the text through the inserted layers instead of handed to it as part of the sequence.

The diagram below draws the sequence for the current bridge. The slider sets how many patch tokens the vision tower contributes; the sequence for the projector and Q-Former designs grows with it, while the cross-attention sequence is flat because the image never enters it. Watch the left-to-right length rather than the contents: for the first two designs every extra patch is another box the language model has to attend over, whereas the cross-attention row stays exactly as long as the text. That difference is the whole argument of the part in one picture.

The language model's input sequence for the selected bridge: instruction text (blue), image or query tokens (pink), and the reply prompt (blue). Drawn with Guide.setupCanvas.

Text tokens are fixed at a handful either side of the image; the image is the variable.

⚠ Prepending image tokens is the default for a reason. It needs no change to the language model at all, so the same recipe works for any frozen checkpoint. The cost is that every image token is a full row of every attention matrix, so the sequence length — and with it the prefill cost and the KV cache — is paid at the input, not hidden in a side channel. Prefill is the first pass in which the model reads the whole prompt before it generates anything, and the KV cache is the running notebook of keys and values it keeps from that pass so it does not have to reread the prompt for every new word; both grow with the prompt length, which is why a long image prefix is expensive.
6

Where this shows up

The bridge names the model

Projectors

LLaVA and the instruction-tuned family

LLaVA pairs a CLIP vision tower with a Vicuna language model and trains a two-layer MLP between them, then fine-tunes on generated visual instructions. The simplicity of the bridge is the reason the recipe reproduces on almost any pair of open checkpoints, which is why there are so many descendants. Because the projector is small and self-contained, swapping in a different vision encoder or a different language model costs mostly the price of retraining that one module, and the rest of the system carries over unchanged.

Resamplers

BLIP-2, InstructBLIP and the query school

BLIP-2 trains a Q-Former between a frozen vision encoder and a frozen language model, and makes the queries the interface. Because the token count is fixed, the same Q-Former can feed different language models, which is exactly the modularity the design was aiming for. Once the Q-Former has turned any image into the same 32 vectors, those vectors are just a short piece of context, and a second language model only has to learn to read them rather than deal with the image at all.

Flamingo's gated cross-attention sits at the other end of the spectrum and is the ancestor of the open interleaved models that followed, where text and images arrive in the same conversation and the model reads them in order. The thing to notice is that all four designs answer the same question — how does one modality's vectors become another's tokens — and that the answer is always a small module between two large frozen ones. The names differ, the parameter counts differ by two orders of magnitude, and the sequence diagrams look nothing alike, but the structural bet is identical: leave the expensive pretrained parts alone and train only the cheap part that lets them cooperate. The token budget that the bridge implies is the subject of the next part.

Further reading

These papers define the four bridges and the training recipes that make them work. Read them together and the design space is small: one projector, one resampler, one fusion layer, and a lot of data. The order below roughly tracks the ideas — a resampler that reads an image with learned queries, a projector that translates patch vectors into word vectors, and an attention layer that feeds image features into a frozen language model — and seeing them side by side makes it clear how little of a vision-language model is actually about vision.

Cheat sheet

TermMeaning here
Bridge / adapterThe trainable module between a frozen vision tower and a frozen language model
Linear projectorOne matrix $d_v \times d_l$; applied per patch, cheapest possible map
MLP projectorA projector with a hidden layer and nonlinearity; the LLaVA choice
Q-FormerA transformer with a fixed set of learned queries that read the image into that many tokens
Perceiver resamplerThe same latent-query idea, used inside Flamingo to shrink a variable patch set
Gated cross-attentionNew attention layers in the language model, gated from zero, so image tokens never enter the sequence
Trainable parametersAlmost always the bridge and nothing else in the first stage; the towers stay frozen
Token budgetSeparate from the parameter budget: set by the patch grid or by the query count
8

Check your understanding

0/4 answered