Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

A language model does not see pixels or sound waves. It sees a sequence of token ids: small integers that stand for chunks of text. Everything it knows comes from letting each position in that sequence attend to the positions before it, which is the operation a transformer performs. The whole problem of multimodal modelling is how to turn a picture — or a second of speech, or a video clip — into something a transformer can attend to, and then how to make the model care about it. Turning the signal into tokens is the easy half: an encoder can chop an image into patches, or a waveform into short frames, and emit one vector per piece. Making the model care is the harder half, because a language model trained only on text has no reason to prefer what is actually in the image over what its language prior expects to see. This volume starts with the simplest possible answer, a grid of patches fed through a transformer, and follows the design space out to the 2026 frontier, where one trunk — a single shared network, rather than a separate vision model bolted onto a language model — consumes text, image, audio and video in the same early-fused token stream and answers in text, speech or pixels.

It is the second half of a pair. Generative Media, Interactively makes media; this volume reads it, and the two meet in the middle at the latent space both share. A latent space is the internal vector space in which a model represents its inputs. Two models share one when similar things end up near each other in it, so a sentence and the image it describes can be compared directly instead of being translated back and forth. The prerequisites are the same foundation volumes — the linear algebra guide for the similarity and projection geometry that contrastive learning is built from (contrastive learning trains a model by pulling matching pairs closer together and pushing mismatched ones apart), and the probability guides for the softmax and the likelihoods behind every caption (softmax turns a row of raw scores into positive numbers that sum to one, which is how a model expresses "how much should I pay attention to each option"). Where it needs the serving side of a multimodal request — the prefill/decode inversion an image token causes, the encoder cache, the disaggregated vision pool — it links straight into LLM Serving, Interactively, whose Part 18 states those costs and whose Part 13 sizes the pools. Those terms are the serving vocabulary for ideas first met here: prefill is the single pass that reads the whole prompt at once, decode is the slower token-by-token generation that follows, and an image reverses their usual balance because it inflates the prompt without extending the answer. The from-scratch encoder it assumes lives in the sibling volume.

The parts

Part 1
Pixels into tokens

Patching an image into a sequence, position embeddings and the class token, and what a vision transformer gives up against a convolution.

Part 2
Caption matches picture

InfoNCE on the similarity matrix, the temperature and batch size it needs, and SigLIP's pairwise sigmoid alternative.

Part 3
What a shared space buys

Zero-shot classification, retrieval and prompt templates, and the compositionality, counting and typographic failures the space inherits.

Part 4
The modality zoo

Audio, video and 3D encoders aligned into one space, what a hub modality binds, and what aligned does and does not mean.

Part 5
Bolting an encoder onto a language model

Linear and MLP projectors, the Q-Former and cross-attention bridges, where image tokens sit, and what each option trains.

Part 6
Resolution and the token budget

Fixed patches, AnyRes tiling and native dynamic resolution, the merge that compresses tokens, and the image's prefill cost.

Part 7
Training a vision-language model

The staged recipe, freezing schedules and data mixtures, what breaks along the way, and preference tuning on multimodal pairs.

Part 8
Why VLMs hallucinate

Objects that were never there, attention sinks on image tokens, the language prior overriding pixels, and the benchmarks that measure it.

Part 9
One trunk, many modalities

Dropping the bridge between towers, native multimodal pretraining, modality embedders into one stream, and modality-aware experts.

Part 10
Video in, long context

Frame sampling against a token budget, temporal position encoding, streaming memory, and the KV arithmetic a clip implies.

Part 11
Listening models

Speech as an encoder's output or as native audio tokens in the trunk, speech instruction following, and paralinguistics.

Part 12
Omni models and the real-time loop

Interleaved token streams, the thinker-and-talker split, listening while speaking, barge-in, and the realtime budget.

Part 13
Understanding and generating

One model that reads and makes images, the tokenizer choice that decides whether it works, and three generation orders.

Part 14
Grounding

Detection and segmentation as coordinate tokens, promptable segmentation, referring expressions, and grounding an interface.

Part 15
Vision-language-action

Actions as text tokens, a flow-matching action expert, chunking, and the latent world models that predict what happens next.

Part 16
Serving and evaluating multimodal systems

Encoder disaggregation and caching, the prefill-and-decode inversion an image causes, safety, and the open problems.

Reference

Start at Part 1 →