Multimodal Models, Interactively
Volume II of the Multimodal & Generative Media arc - how a model is made to see, how a vision encoder is bolted onto a language model, and how one trunk comes to read text, images, audio and video together, then how it grounds, acts and gets served.
A language model does not see pixels or sound waves. It sees a sequence of token ids: small integers that stand for chunks of text. Everything it knows comes from letting each position in that sequence attend to the positions before it, which is the operation a transformer performs. The whole problem of multimodal modelling is how to turn a picture — or a second of speech, or a video clip — into something a transformer can attend to, and then how to make the model care about it. Turning the signal into tokens is the easy half: an encoder can chop an image into patches, or a waveform into short frames, and emit one vector per piece. Making the model care is the harder half, because a language model trained only on text has no reason to prefer what is actually in the image over what its language prior expects to see. This volume starts with the simplest possible answer, a grid of patches fed through a transformer, and follows the design space out to the 2026 frontier, where one trunk — a single shared network, rather than a separate vision model bolted onto a language model — consumes text, image, audio and video in the same early-fused token stream and answers in text, speech or pixels.
It is the second half of a pair. Generative Media, Interactively makes media; this volume reads it, and the two meet in the middle at the latent space both share. A latent space is the internal vector space in which a model represents its inputs. Two models share one when similar things end up near each other in it, so a sentence and the image it describes can be compared directly instead of being translated back and forth. The prerequisites are the same foundation volumes — the linear algebra guide for the similarity and projection geometry that contrastive learning is built from (contrastive learning trains a model by pulling matching pairs closer together and pushing mismatched ones apart), and the probability guides for the softmax and the likelihoods behind every caption (softmax turns a row of raw scores into positive numbers that sum to one, which is how a model expresses "how much should I pay attention to each option"). Where it needs the serving side of a multimodal request — the prefill/decode inversion an image token causes, the encoder cache, the disaggregated vision pool — it links straight into LLM Serving, Interactively, whose Part 18 states those costs and whose Part 13 sizes the pools. Those terms are the serving vocabulary for ideas first met here: prefill is the single pass that reads the whole prompt at once, decode is the slower token-by-token generation that follows, and an image reverses their usual balance because it inflates the prompt without extending the answer. The from-scratch encoder it assumes lives in the sibling volume.
The parts
Patching an image into a sequence, position embeddings and the class token, and what a vision transformer gives up against a convolution.
InfoNCE on the similarity matrix, the temperature and batch size it needs, and SigLIP's pairwise sigmoid alternative.
Zero-shot classification, retrieval and prompt templates, and the compositionality, counting and typographic failures the space inherits.
Audio, video and 3D encoders aligned into one space, what a hub modality binds, and what aligned does and does not mean.
Linear and MLP projectors, the Q-Former and cross-attention bridges, where image tokens sit, and what each option trains.
Fixed patches, AnyRes tiling and native dynamic resolution, the merge that compresses tokens, and the image's prefill cost.
The staged recipe, freezing schedules and data mixtures, what breaks along the way, and preference tuning on multimodal pairs.
Objects that were never there, attention sinks on image tokens, the language prior overriding pixels, and the benchmarks that measure it.
Dropping the bridge between towers, native multimodal pretraining, modality embedders into one stream, and modality-aware experts.
Frame sampling against a token budget, temporal position encoding, streaming memory, and the KV arithmetic a clip implies.
Speech as an encoder's output or as native audio tokens in the trunk, speech instruction following, and paralinguistics.
Interleaved token streams, the thinker-and-talker split, listening while speaking, barge-in, and the realtime budget.
One model that reads and makes images, the tokenizer choice that decides whether it works, and three generation orders.
Detection and segmentation as coordinate tokens, promptable segmentation, referring expressions, and grounding an interface.
Actions as text tokens, a flow-matching action expert, chunking, and the latent world models that predict what happens next.
Encoder disaggregation and caching, the prefill-and-decode inversion an image causes, safety, and the open problems.