One trunk, many modalities
Every model so far in this volume has the same silhouette: a strong encoder on one side, a language model on the other, and a thin bridge in between. To keep the vocabulary straight: a bridge is the small module that translates one tower's vectors into the other's format — a projector, a resampler, or a few cross-attention layers, depending on the design. That bridge exists because the encoder was trained first, on its own, and the language model was trained first, on text, so neither was ever asked to share a sequence. Early fusion — the name for mixing the modalities at the input, before any deep processing, rather than encoding each one separately and comparing them afterwards — deletes the bridge by refusing that premise. Instead of building two specialists and teaching them to talk, you train one trunk, a single shared backbone network, from scratch on one interleaved stream of text, image and audio tokens, meaning a sequence in which the tokens of every modality appear in the order they actually occurred: a sentence, then the image or sound that accompanied it, then the next sentence. Each modality gets its own embedder at the front of that trunk, a small front-end network that turns the modality's raw tokens into vectors the shared trunk can read, and the same attention then does all of the mixing. You can picture the difference as one shared language learned from birth, rather than two adults raised apart who need a translator standing between them. This part is about what the bridge costs, what replaces it, and where the replacement is still hard.
The bridge and its cost
Two towers, one interface
A vision-language model is usually assembled from parts rather than trained as a single whole. A vision encoder is trained with a contrastive or self-supervised objective — both are ways of teaching it to represent images without any help from a language model — and its weights are then frozen, meaning they are held fixed and never updated again. A small projector then maps its image vectors into the language model's embedding width, where width is how many numbers each vector contains; the projector has to produce vectors of exactly the width the language model expects, because that model will only accept its own input channel. The language model then reads image tokens as if they were foreign words. This arrangement is what the field often calls late fusion: the two modalities are encoded separately, in their own towers, and only meet at the very end, where a translator stands between them. This is the recipe in the adapter part, and it works marvellously — it is also a chain of compromises. The first compromise is that the encoder never learns what the language model needs, because it was trained before the language model entered the picture and is frozen by the time the two are joined. The second is that the projector is the only interface, and therefore the only thing that trains, so all of the model's cross-modal skill has to be squeezed through that one small module. The third is that the image features are frozen at whatever the encoder happened to find important, whether or not those happen to be the features a question about the image would need.
Early fusion throws the interface away. One transformer trunk is pretrained on a stream of interleaved tokens from every modality, with a per-modality embedder at the input and a per-modality head at the output. Pretrained here means trained before the model is specialised for any particular task, on a very large stream of data. The training signal is next-token prediction: the model is repeatedly shown a prefix of the stream and asked to guess the token that comes next, whichever modality that token belongs to. An embedder is the module that turns one modality's raw tokens into vectors of the trunk's width, and a head is the matching module at the far end that turns the trunk's output back into whatever the task needs, such as the next token to emit. Attention runs across the whole mixed sequence, so a text token can attend to an audio token directly, with no projection in between and no frozen tower to work around. That matters because the mixing is not a one-off step at the entrance: it happens again at every layer, so the model can keep refining how the modalities inform one another as the representation deepens. The diagram contrasts the two arrangements; the slider picks which one is lit.
Left: a bridged two-tower model, with a projector standing between a frozen encoder and the language model. Right: early fusion, where each modality enters through its own embedder into one shared trunk.
One trunk, one context
A shared budget and attention that crosses modalities
In an early-fused model there is a single context window, and every token competes for it. The context window is the fixed maximum number of tokens the model can hold in mind at once; anything beyond it has to be dropped or summarised. The competing tokens are subwords (the word pieces a text tokeniser produces), image patches (the small tiles an image is cut into), audio frames (short slices of sound), and video latents (compressed per-frame features). That is the first real consequence of fusing early: the modalities are not in separate streams joined at the end, they are in one budget that has to be divided. Concretely, if the window holds twelve thousand tokens, a long video does not get "as much room as it needs"; it gets whatever is left after the text and the image tokens, and the division is a design decision with a visible cost. Spend that budget on a fine-grained image and the audio has less room to describe what was said, and the other way round.
The second consequence is what attention does with a mixed sequence. Attention is the operation that lets each token gather information from other tokens by weighing how relevant they are to it, and it is deliberately blind to what kind of token it is looking at; nothing in the mechanism says "this is an audio frame, do not read it". Because every token can attend to every other, the attention matrix is not block-diagonal — that is, the strong weights are not confined to one square block per modality, with the gaps between blocks left empty. Instead, text tokens attend to image patches, audio frames attend to neighbouring speech and to the text around them, and image regions attend to the text that describes them. Each of those is a cross-modal connection, one that crosses from one modality to another. The heatmap below builds such a matrix for the token mix on the sliders, blocking it by modality so you can see the diagonal blocks (within-modality attention) and the off-diagonal ones (cross-modal attention). Turn a modality's budget down and its block shrinks; the cross terms stay, which is the whole point of fusing early — the connections between modalities survive even as a modality's share of the budget changes.
Qwen3-VL's DeepStack is a concrete instance of that claim. Instead of injecting the vision encoder's output only at the input, where the patches are handed to the trunk and then left behind, DeepStack taps the vision tower at several depths and adds those intermediate features into the matching layers of the language trunk. Early layers of the vision tower carry fine, local detail and later ones carry coarser, more semantic summaries, so routing each level to a different depth of the trunk gives the language model access to the whole hierarchy rather than a single final vector. The effect is that vision and language mix repeatedly rather than once at the door: information crosses between the modalities at many points along the stack, and that is what mixing at every layer means in practice when the vision tower was trained separately and only the trunk is shared.
The shared context budget, drawn with Guide.drawStacked. The dashed marker is the window capacity, and the three segments are the text, image and audio token counts on the sliders.
A mixed-modality attention matrix from Guide.drawHeatmap: rows are query tokens, columns are keys, labelled T for text, I for image and A for audio. Within-modality blocks are bright by construction; the off-diagonal blocks are the cross-modal attention early fusion makes possible.
Modality embedders
Different front ends, one width
The trunk cannot be fed raw patches and raw waveform frames any more than it can be fed raw text. A tokeniser is the rule that decides what counts as one token for a given modality: it chops a continuous signal into discrete units, and the trunk only ever sees those units. Each modality gets an embedder — the direct successor of the projector in the bridged design, but trained jointly with the trunk rather than bolted on after. That difference matters: the projector was taught to translate into a frozen language model that could not adapt, whereas the embedder and the trunk are optimised together, so the trunk can learn to make sense of whatever the embedder learns to send. For text and images the embedder is close to the familiar one: a learned lookup for subwords (an embedding table, one row of numbers per subword, fetched by index) and a patch projection for images (a small matrix applied to each patch vector). Audio is the interesting case, because it can be embedded either as encoder features or as discrete codec tokens, which is the subject of a later part. Either way the embedder's job is the same: take whatever the modality's tokeniser emits and produce vectors of the trunk's width.
The embedders differ wildly in how many tokens they contribute per unit of signal, and that rate is what the budget in the previous step is dividing. Forty words of text are roughly forty tokens, but a single modest image can contribute hundreds, one second of codec audio can contribute hundreds more, and a short clip multiplies again by its frame count. The bars put a short sentence, a single image, one second of codec audio and a short clip on the same axis so the mismatch is visible: the same "one unit" of each modality is a different order of magnitude of sequence length. That mismatch is why the budget is a real constraint rather than a formality. Because every one of those tokens becomes a row of the attention matrix, token count drives both the compute and the memory, so the modality with the most generous embedder quietly decides how much of the window the others are allowed.
Tokens contributed by one unit of each modality: 40 words of text, one 336-pixel image at patch P, one second of codec audio at 75 Hz × 8 books, and an 8-frame clip. Image tokens come from Multimodal.cost.imageTokens, audio from GenMedia.dsp.tokenRate and video from GenMedia.cost.videoTokens.
A 336 px image at patch 14 is 576 tokens; the same image at patch 8 is 1764. The embedder's front end, not the trunk, decides that number.
Modality-aware experts
Shared attention, specialised feed-forward
One trunk does not have to mean one set of weights everywhere. Every transformer block separates two jobs: attention, which mixes information across positions, from the feed-forward layer, which transforms each position on its own, with no reference to its neighbours. A mixture-of-experts (MoE) trunk exploits that split. It keeps attention shared — attention is what lets text look at an image, so it must be common to every modality — and replaces the single feed-forward layer with a router and a pool of experts. An expert is an ordinary feed-forward layer with its own parameters, and the router is a small learned function that looks at a token and decides which experts should process it. Each token is sent to a few experts rather than all of them, and the router is free to learn that some experts are better at patch-like tokens and others at subwords or audio frames. The payoff is capacity: capacity meaning how much stored knowledge the model can hold, which here grows with the size of the expert pool even though each token only ever uses a few experts.
That is the "modality-aware" part, and it is a soft specialisation, not a hard rule: nothing forces expert 6 to serve audio, but a router trained on a mixed stream tends to discover the split, because the statistics of image patches and speech frames really are different. Keeping the split soft is the point: a hard rule would have to be written down by the designer in advance, and would break the moment a new modality arrived, whereas a learned router can be renegotiated by training. The heatmap shows routing weights for twelve tokens — four from each modality — into eight experts, with the top-k slider deciding how many experts each token may use. Here "top-k" means the router keeps only the k highest-scoring experts for each token and ignores the rest. A low k is a hard, sparse route, so each token touches few parameters and the model is cheap per token; a high k is nearly a dense layer with extra parameters, so more of the pool contributes and the compute approaches that of an ordinary feed-forward layer.
Router weights from a seeded stream, kept only for each token's top-k experts and renormalised. Rows are tokens (T, I or A by modality), columns are experts; the bright columns per group are the specialisation the router discovers.
Sparsity buys parameters without paying for them at every token: the trunk's compute stays close to a dense model of the same active width, while the total parameter count grows with the expert pool.
Where this shows up
Fused trunks in the wild
Mixed-modal pretraining from scratch
The Chameleon family trains a single token-based trunk on interleaved text and images, with a shared vocabulary and no separate vision tower. A shared vocabulary means the same discrete token set covers image patches and text subwords, so the trunk does not even have to know which modality a given token came from. Gemini is described the same way at scale: one model that ingests interleaved input. This is what the cheat sheet calls native multimodal pretraining — training the trunk on mixed modalities from the start, rather than adapting a model that learned text first. The price is that the vision front end has to be pretrained jointly, which is exactly the step the bridged recipe avoids: you cannot start from a vision encoder that already knows images and a language model that already knows text, because there is only one model and it has to learn both at once.
When the pretrained encoder is the point
The bridge is not obsolete. It lets a team reuse a SigLIP or DINOv2 encoder that took a large compute budget to train, and it lets the language model stay untouched, so a project can stand on two pieces of work that other people already paid for. Flamingo's gated cross-attention and BLIP-2's Q-Former are bridges by another name, and they remain the pragmatic choice when the vision tower is better than anything you could pretrain jointly. That last condition is the deciding one: if a standalone vision encoder trained on an enormous image corpus sees better than a trunk you could afford to train on mixed data, then freezing it and training a translator is the better use of your budget.
The trade is legible once you name the two costs. A bridge trains a small interface on top of frozen parts, so its cost is the price of the bridge itself and its ceiling is whatever the frozen parts already know. Early fusion trains one large model on interleaved data, so its cost is a full pretraining run and its reward is attention that crosses modalities directly. Neither option dominates: the choice follows the encoders you already have and the compute you can spend. The next parts push the fused trunk into the two modalities that make it hardest, video with its length and audio with its continuity, then assemble the real-time omni model that uses both.
Further reading
The bridge literature and the early-fusion literature are two answers to the same question about where modality mixing should happen: at the end, through a translator between frozen towers, or from the start, inside one trunk. These papers give the bridge in its two best-known forms, the fused trunk at scale, and the expert routing that a shared trunk can adopt. Read in order, they move from reusing pretrained parts to training everything together, which is the arc of this part.
- Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc and coauthors, "Flamingo: a Visual Language Model for Few-Shot Learning", 2022 — frozen vision and language towers joined by gated cross-attention, a bridge with extra depth.
- Junnan Li, Dongxu Li, Silvio Savarese and Steven Hoi, "BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models", 2023 — the Q-Former, a lightweight querying bridge between frozen halves.
- Chameleon Team, "Chameleon: Mixed-Modal Early-Fusion Foundation Models", 2024 — one token-based trunk over interleaved text and images, with no separate vision encoder.
- Gemini Team, "Gemini: A Family of Highly Capable Multimodal Models", 2023 — a natively multimodal model family described as trained on interleaved input from the start.
- Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux and coauthors, "Mixtral of Experts", 2024 — sparse mixture-of-experts routing, the mechanism behind modality-aware feed-forward layers.
- Shuai Bai, Yuxuan Cai, Ruizhe Chen and coauthors, "Qwen3-VL Technical Report", 2025 — DeepStack multi-level vision injection.
Cheat sheet
| Term | Meaning here |
|---|---|
| Bridge | The projector, resampler or cross-attention layers joining a frozen encoder to a language model |
| Early fusion | One trunk pretrained on interleaved tokens from every modality, with no separate towers |
| Modality embedder | The per-modality front end that maps a tokeniser's output into the trunk's width |
| Shared context | One window that all modalities compete for; the division is a fixed design decision |
| Cross-modal attention | Attention weights off the modality diagonal: text attending to patches, speech to text |
| Modality-aware experts | A shared attention stack with a routed feed-forward pool that can specialise per modality |
| Native pretraining | Training the trunk on interleaved modalities from the start, rather than adapting a text model |
| What fusion does not fix | Token budget, embedder cost and the quality of the modality tokeniser all remain |