Bolting an encoder onto a language model
A vision encoder produces a stack of patch vectors; a language model expects token embeddings of its own width. To keep the vocabulary straight: a tower is one of the two large sub-networks in the pair — here the vision encoder that reads pixels and the language model that reads and writes words — and each tower was trained on its own kind of data, so each has its own idea of what a vector should look like. A language model, for instance, only knows how to consume vectors of one particular width, the size of its own embedding space, and it has no reason to expect a patch vector to already have that width, that scale, or that internal arrangement of meaning. Nothing about either tower guarantees that these two things live in the same vector space, and no amount of contrastive pretraining forces the pair to line up dimension by dimension. A bridge has to sit between them. It is a small, deliberately unglamorous piece of engineering, and it is where almost every trainable parameter in a vision-language model lives. This part is about the four bridge designs you will meet, what each of them trains, and where the image ends up in the token sequence.
The interface between towers
Two frozen spaces and one learned adapter
A vision-language model starts from two pretrained parts that were trained separately. Pretrained means someone already spent the compute and the data to train each part on its own task — one to recognise images, one to predict text — and we get to reuse the result instead of starting over. The vision tower is a patch transformer — the encoder of the previous parts — and it emits one vector per patch at a width of, say, $1024$. Width here means how many numbers are packed into each vector. The language model expects vectors of width $4096$, and it will only accept them as the same input channel it uses for text. The bridge is whatever maps one to the other, and it has to do so for a whole grid of patch vectors at once, because the vision tower hands over every patch rather than one summary of the image. Here is the important structural choice, and it is worth pausing on because it explains almost everything else in this part: the two towers are usually frozen, or nearly so, during training, and only the bridge between them is updated. Frozen means a tower's weights are held fixed — they are read but never changed — so no training signal flows back into it. Why does that matter? Training a vision encoder from scratch, or a language model from scratch, takes enormous compute and enormous data — that is exactly why pretrained versions of each already exist and are reused here. If you unfroze both towers and trained everything jointly, you would be re-paying that entire cost just to teach the pair to talk to each other. Freezing them instead means you only ever have to train the (comparatively tiny) bridge — a translator that has to learn to speak two already-fluent languages to each other, not learn either language from scratch. That is the single fact that makes vision-language training affordable at all.
Four designs account for nearly everything in the literature. A linear projector is a single learned matrix, cheap and surprisingly strong; it is called linear because all it can do is a weighted sum, with no bend or curve in the mapping. An MLP projector adds a hidden layer and a nonlinearity, which is what LLaVA settled on. A hidden layer is an extra stage of computation between the input vector and the output vector, and a nonlinearity is a function that lets that stage bend its input rather than only rescale it — without one, stacking layers would collapse back into a single matrix. A Q-Former, from BLIP-2, uses a small transformer with a fixed set of learned query tokens that read from the image, which compresses hundreds of patches into a handful of vectors; a query token here is a learned vector that asks a question of the image features, much as a search term asks a question of a database. Gated cross-attention, from Flamingo, inserts attention layers into the language model itself so the text can look at the image without any image tokens entering the sequence at all. Cross-attention is attention in which the queries come from one stream, here the text, and the keys and values come from another, here the image, so one stream can read the other without the two being merged into a single sequence. The gate is a learned multiplier that starts at zero, which means each inserted layer is switched off at the beginning and only turns on as training finds it useful.
Everything else in this part hangs off that menu. The demo below prices the bridges: pick a design and a vision tower, and watch the trainable parameter count move. Parameter count is a rough proxy for two things at once, how much memory the bridge needs and how much data it will take to train, which is why it is the number people usually quote first. The gap between the cheapest and the most expensive option is two orders of magnitude, and that gap is the reason the choices are made the way they are — a team with a modest compute budget and a team with a large one will not reach the same conclusion from the same menu, and both can be right for their setting.
Projectors, linear and otherwise
The smallest thing that could possibly work
The linear projector is a single matrix $W$ that takes a vision vector of width $d_v$ to a language vector of width $d_l$. It has $d_v \cdot d_l$ parameters, it is trained on image-caption or instruction data, and it is applied to every patch token independently. That is the whole idea. What makes it work is that the vision tower has already been aligned to language by a contrastive objective, so the geometry is roughly right and the matrix only has to rotate and rescale it. In other words, the contrastive stage already pushed image vectors and word vectors into comparable positions, so the projector's job is closer to translation than to re-invention: it moves the vectors to the right place without having to decide afresh what they mean. Adding a hidden layer — an MLP with a nonlinearity, the LLaVA design — costs another $d_l^2$ parameters and buys a slightly more expressive map. A hidden layer is an intermediate vector the network computes before the output, and a nonlinearity is a bending or squashing function applied to it; together they let the projector fit a curved mapping rather than only a straight one, which is the sense in which it is more expressive. The MLP is worth its cost at $d_l = 4096$ and is not worth much more than that; three-layer projectors have been tried and rarely win.
Each of these options is applied per token, so the number of trainable parameters does not depend on how many patches the image produces. That is the property to hold on to: a projector's parameter count is set by the two widths, while its token count is set by the patch grid. The two budgets are separate levers. A taller image or a finer patch size will hand the language model more rows to read, but it will not add a single trainable weight to the projector, because the same small matrix is reused on every row. This is why the rest of the part can discuss a parameter budget and a token budget almost independently: one is fixed by the widths you chose, the other by the resolution you feed in.
Trainable parameters for each bridge at the selected vision tower, drawn with Guide.drawBars. The highlighted bar is the current choice; the language model itself is assumed frozen.
Estimates for a 4096-wide language model and a 768-wide Q-Former; the cross-attention bar counts four inserted gated layers. Treat them as orders of magnitude, not exact counts.
Querying the image with a Q-Former
A fixed number of queries, however big the image
BLIP-2's Q-Former is a small transformer with a fixed set of learned query tokens — typically $32$ of them — that cross-attend to the vision tower's patch features. Cross-attention means the queries come from one place and the keys and values from another: here the queries are the learned tokens and the keys and values are the patches. The queries do not depend on the image; they are parameters. What depends on the image is the attention, so each query ends up pulling a different weighted average of the patches, and the $32$ output vectors become the $32$ image tokens the language model sees. Picture a panel of $32$ specialists, each with a fixed interest, all reading the same page of patches: one tracks colour, another tracks shapes on the left, and each reports the average of whatever caught its attention. The trick is that the number $32$ is chosen by the designer, not by the resolution: a Q-Former turns an image of any size into the same short summary, which keeps the language model's context budget predictable and makes the whole system resolution-agnostic at the interface. The context budget is the number of tokens the language model can hold at once; because it is fixed, a bridge that always emits the same number of image tokens fits a fixed budget no matter how large the picture is.
The cost is that a summary is lossy. Thirty-two vectors have to carry everything the language model will ever know about the picture, and fine spatial detail — the serial number on a label, the position of a small object — is exactly what a soft average of patches destroys. The reason is that averaging is a smoothing operation: if one patch holds a tiny, unique mark and its neighbours hold background, the weighted average keeps the background and washes the mark out. The heatmap below shows the query-to-patch attention: each row is a query, each column a patch, and a sharp row is a query that has found something specific to report. The temperature slider controls how sharply each query commits; a low temperature concentrates a row onto a few patches, and a high temperature spreads it across the whole image.
Query-to-patch attention, softmaxed per row. Rows are the 32 learned queries, columns are image patches; the matrix comes from a seeded stream, so it is the same on every load.
A fixed query set means a fixed token count. That is the point of the design and also its ceiling.
Cross-attention instead
Leave the language sequence alone
Flamingo took a different route. Rather than converting patches into tokens that sit in the language model's input, it inserted new cross-attention layers into the frozen language model's blocks, so that the text tokens can attend to the image features at several depths. There are no image tokens in the sequence at all: the text is generated exactly as before, except that every so often it is allowed to look at a set of image features that never enter the self-attention. Each inserted layer carries a learnable gate initialised at zero, so training starts from the frozen model's behaviour and gradually opens the tap — a stabiliser that matters when the rest of the network is frozen and the new layers see gradients for the first time. Gradients are the signals training uses to nudge each weight; a fresh, randomly initialised layer produces large and noisy ones, and if those signals could flow back into the pretrained text model they would disturb abilities that already work. Starting the gate at zero means the new layer contributes nothing at first, so the model's outputs are unchanged and the gradients stay small until the gate has learned something worth adding.
The advantage is that the language model's sequence length is untouched, so nothing in its positional or context story has to change, and the frozen text ability is protected. That matters because a language model has a fixed maximum sequence length and a fixed sense of where each position sits; adding hundreds of image tokens pushes against both, and a bridge that inserts nothing avoids the problem entirely. The disadvantage is that you now have trainable layers inside the language model rather than beside it, which is a larger and more intrusive edit: you are modifying the model's own blocks, so the bridge has to be built for that specific architecture rather than attached from outside. The heatmap below is text-token attention over image tokens with the gate applied; drag the gate down and the image contribution fades toward the frozen text-only model.
Effective cross-attention of text tokens (rows) over image tokens (columns), scaled by the gate. At gate zero the inserted layers are transparent and the model falls back to its text-only behaviour.
The gate starts at zero and is learned. It is the standard answer to "how do I add a new pathway to a frozen model without wrecking it".
Where the image tokens go
One prefix, one bridge, no free lunch
Once a bridge is chosen, the resulting image representation has to be spliced into the language model's input, and where it lands is a design decision with consequences. LLaVA concatenates the projected image tokens into the sequence before the instruction text, so the language model sees a row of vectors that stand for the picture followed by the words. That chunk of image vectors sitting at the front of the input is called a prefix, and the model has no way to tell it apart from ordinary text except by where it sits and what it contains; the words come after it and attend to it like any other token. The image tokens then attend to the text and the text to the image, both through the ordinary self-attention of the frozen language model, and the number of extra positions is exactly the number of image tokens. A Q-Former inserts its fixed query set at the same place, in the same way, just far fewer of them. Gated cross-attention inserts nothing: the image is a side channel, delivered to the text through the inserted layers instead of handed to it as part of the sequence.
The diagram below draws the sequence for the current bridge. The slider sets how many patch tokens the vision tower contributes; the sequence for the projector and Q-Former designs grows with it, while the cross-attention sequence is flat because the image never enters it. Watch the left-to-right length rather than the contents: for the first two designs every extra patch is another box the language model has to attend over, whereas the cross-attention row stays exactly as long as the text. That difference is the whole argument of the part in one picture.
The language model's input sequence for the selected bridge: instruction text (blue), image or query tokens (pink), and the reply prompt (blue). Drawn with Guide.setupCanvas.
Text tokens are fixed at a handful either side of the image; the image is the variable.
Where this shows up
The bridge names the model
LLaVA and the instruction-tuned family
LLaVA pairs a CLIP vision tower with a Vicuna language model and trains a two-layer MLP between them, then fine-tunes on generated visual instructions. The simplicity of the bridge is the reason the recipe reproduces on almost any pair of open checkpoints, which is why there are so many descendants. Because the projector is small and self-contained, swapping in a different vision encoder or a different language model costs mostly the price of retraining that one module, and the rest of the system carries over unchanged.
BLIP-2, InstructBLIP and the query school
BLIP-2 trains a Q-Former between a frozen vision encoder and a frozen language model, and makes the queries the interface. Because the token count is fixed, the same Q-Former can feed different language models, which is exactly the modularity the design was aiming for. Once the Q-Former has turned any image into the same 32 vectors, those vectors are just a short piece of context, and a second language model only has to learn to read them rather than deal with the image at all.
Flamingo's gated cross-attention sits at the other end of the spectrum and is the ancestor of the open interleaved models that followed, where text and images arrive in the same conversation and the model reads them in order. The thing to notice is that all four designs answer the same question — how does one modality's vectors become another's tokens — and that the answer is always a small module between two large frozen ones. The names differ, the parameter counts differ by two orders of magnitude, and the sequence diagrams look nothing alike, but the structural bet is identical: leave the expensive pretrained parts alone and train only the cheap part that lets them cooperate. The token budget that the bridge implies is the subject of the next part.
Further reading
These papers define the four bridges and the training recipes that make them work. Read them together and the design space is small: one projector, one resampler, one fusion layer, and a lot of data. The order below roughly tracks the ideas — a resampler that reads an image with learned queries, a projector that translates patch vectors into word vectors, and an attention layer that feeds image features into a frozen language model — and seeing them side by side makes it clear how little of a vision-language model is actually about vision.
- Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc and coauthors, "Flamingo: a Visual Language Model for Few-Shot Learning", 2022 — the perceiver resampler and gated cross-attention layers inserted into a frozen language model.
- Junnan Li, Dongxu Li, Silvio Savarese and Steven Hoi, "BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models", 2023 — the Q-Former and its fixed query set.
- Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Jae Lee, "Visual Instruction Tuning", 2023 — LLaVA, the MLP projector and machine-generated multimodal instructions.
- Andrew Jaegle, Felix Gimeno, Andrew Brock and coauthors, "Perceiver: General Perception with Iterative Attention", 2021 — the latent-query resampling that both the Q-Former and Flamingo's resampler inherit.
Cheat sheet
| Term | Meaning here |
|---|---|
| Bridge / adapter | The trainable module between a frozen vision tower and a frozen language model |
| Linear projector | One matrix $d_v \times d_l$; applied per patch, cheapest possible map |
| MLP projector | A projector with a hidden layer and nonlinearity; the LLaVA choice |
| Q-Former | A transformer with a fixed set of learned queries that read the image into that many tokens |
| Perceiver resampler | The same latent-query idea, used inside Flamingo to shrink a variable patch set |
| Gated cross-attention | New attention layers in the language model, gated from zero, so image tokens never enter the sequence |
| Trainable parameters | Almost always the bridge and nothing else in the first stage; the towers stay frozen |
| Token budget | Separate from the parameter budget: set by the patch grid or by the query count |