Understanding and generating
The models in this volume so far read a picture and answer in words. A unified model does the same thing in the other direction: it takes a text prompt and produces an image, with the same trunk that answers questions about one. The trunk here is the shared backbone network that does the actual sequence processing — one set of weights is asked to handle both jobs. That sounds like a small extension and it is not. Reading is a classification problem — every position has one right next token and cross-entropy grades it, meaning the standard loss that rewards putting probability on the correct option. Making is a sampling problem — the model has to choose a plausible whole from a distribution it has only ever graded pointwise, one position at a time. The asymmetry is easy to miss: reading always has a ground-truth answer to compare a prediction against, while a generated image has no single correct target to score, only a distribution to draw from. The tokenizer in between — the rule that decides what counts as one token — decides whether the two objectives can even share a representation, and the three generation orders in use, meaning the sequences in which image tokens are produced, decide what the model conditions on while it samples. This part builds the unified model token by token, then shows where the two objectives collide.
One model for both
The same trunk, two directions
A vision-language model maps pixels to words through a bridge — the small module that translates one tower's vectors into the other's format — and a language model. A unified model keeps one trunk and asks it to do the mapping both ways: image tokens in, text tokens out for understanding, which here means reading a picture and answering questions about it; text plus image tokens in, more image tokens out for generation, which means producing a new picture rather than describing an existing one. Picture a single model that can caption a photograph and then, handed the caption "a red mug on a table", draw a fresh picture of one. The attraction is not that the architecture is elegant — it is that understanding and generation are complementary training signals for the same representation. They exercise different skills on the same underlying content, so together they push the trunk to learn more than either could alone. A model that must draw an object has to learn how its parts fit together and how it looks from different angles; a model that only answers questions about it does not have to learn any of that. Think of the two tasks as two ways of studying the same object, describing it and reconstructing it from memory, where the second exposes structure the first can skate past.
What makes this hard is that the two directions want different things from the token stream, the sequence of tokens the model reads and writes. Understanding wants a token whose distribution can be made sharp, because cross-entropy rewards confidence: it gives a low loss only when the model puts most of its probability on the one right answer. Generation wants a token that is easy to sample from and whose neighbours are visually plausible, because here there is no single right answer and a single wrong jump is a visible artefact — a smudge, a broken edge, an object that suddenly changes identity. Those two preferences do not point the same way. When the tokens are discrete and learned, so that each image token is one of a finite set of categories the model learned, the two pull on the same quantizer, the component that decides which category each piece of image becomes; when the trunk predicts continuous latents with a diffusion head, it stops predicting a category at all. A quantizer tuned for confidence would rather make the gap between neighbouring codes enormous, so the right id is never in doubt; a generator would rather the codes it samples blend smoothly into nearby ones, so a small mistake is not visible. The rest of the part is about that tension, starting with the tokenizer at the front.
The tokenizer crux
A learned codebook can collapse; a fixed grid cannot
An autoencoder turns an image into a grid of continuous vectors — an encoder compresses the picture into those vectors, and a decoder reconstructs the picture back out of them — but a language-model trunk wants discrete ids, whole integers from a countable vocabulary. The classic bridge between those two worlds is a vector-quantized variational autoencoder. Think of tokenizing an image as cutting a photo into a contact sheet of equal tiles and reading them left to right, top to bottom, like words. Vector quantisation is the rule that gives each tile one word: the encoder vector for a tile is snapped to its nearest entry in a learned codebook, a fixed-size table of vectors that acts as the vocabulary, and the decoder reconstructs from those entries rather than from the original vectors. Training adds a commitment loss that pulls encoder outputs toward codebook entries and entries toward encoder outputs, so the two meet in the middle. Concretely, two tiles of sky that look similar have encoder vectors that sit close together, and both may snap to the same entry; the table is what turns an unbounded set of vectors into a countable set of words. The failure mode is exactly what that coupling invites. Early in training a few entries are slightly closer, they win more vectors, they get pulled closer still, and the rest of the codebook becomes dead weight. The model still reconstructs, using a handful of entries, and the discrete bottleneck — the narrow channel of discrete choices every tile must pass through — has quietly become a coarse quantizer with most of its capacity unused. That is codebook collapse, and it degrades generation most: sampling has fewer distinct tokens to compose with, so the model has a smaller vocabulary for building an image than it was supposed to.
FSQ — finite scalar quantization — removes the thing that collapses. Instead of a learned table, each coordinate of the latent is rounded against a fixed set of levels. If a tile's latent vector has three coordinates and each is rounded to one of four levels, the code is just the tuple of those three per-coordinate indices, and there are 4 × 4 × 4 = 64 possible codes. There is no nearest-neighbour search, no commitment loss and no table to starve: every level is reachable by construction, and the implicit vocabulary size is the product of the level counts. The codebook still exists as a mathematical set of codes, but it is implied by the grid rather than learned, so nothing in training can pull usage toward a few favoured entries. The trade is that a fixed grid cannot adapt to the data the way a learned table can: it is safe, but it never discovers entries that the images would have rewarded. The bars below show what happens to a codebook's usage as an auxiliary pressure is turned up. The learned codebook concentrates onto a few entries; the fixed grid stays close to uniform, because nothing in its objective can reward concentration.
Codebook entry usage for a learned VQ codebook and for a fixed FSQ grid, normalised by the busiest entry. The collapse slider raises the pressure that drives a VQ table to concentrate; the FSQ bars do not move.
Three generation orders
What the model conditions on while it decides
A raster (row-major) order reads image tokens like a sentence: top-left to bottom-right, one token at a time, each conditioned on everything before it. Conditioned on means the model's prediction for the current token is allowed to look at every token already produced. It is the order a plain autoregressive transformer already understands — autoregressive is the standard recipe of predicting the next token from the ones before it, one step at a time — and it is the weakest for images, because a token's nearest neighbours in space are far away in the generation order. The token below a pixel is generated hundreds of steps after the token above it, so the model has to remember a lattice in a one-dimensional stream. By the time it reaches the bottom row, the top row is far behind in the sequence and the spatial structure it needs is no longer nearby. Raster is nevertheless the order a text model can adopt with no architectural change, which is exactly why it was tried first: the sequence is already one-dimensional, and next-token prediction needs no modification at all.
VAR — visual autoregressive modelling — changes the axis rather than the mechanism. It generates the image as a sequence of scales, meaning resolutions: a single coarse token for the whole picture first, then a $2 \times 2$ grid, then $4 \times 4$, each scale conditioned on all coarser scales. Picture an artist blocking in a painting: first the rough shape and colour of the whole scene, then progressively finer strokes, each consistent with the block-in underneath. Because the next scale is predicted in parallel, with all the tokens of a scale produced at once rather than one at a time, the number of sequential steps is logarithmic in resolution instead of quadratic, and every token at a scale sees the global structure before it commits to detail. In other words, doubling the resolution adds a step rather than quadrupling the work. For the four scales 1×1 through 8×8, that is four sequential passes rather than eighty-five token steps. MAR — masked autoregressive — takes the opposite route to the same goal. It reveals positions in a seeded order, a pseudo-random sequence fixed by a seed so it can be reproduced, and at each step predicts a whole set of masked positions at once, conditioned on all revealed ones. Because the revealed tokens can sit on every side of a masked one, the model can use context on both sides of a token, not just the tokens before it. The buttons switch the reveal order on a small grid and animate it; the picture is the same, only the order in which it is committed changes.
A $32 \times 32$ test image drawn with GenMedia.raster and revealed as an $8 \times 8$ grid of cells. The order comes from Multimodal.gen.rasterCells, Multimodal.gen.nextScaleCells and Multimodal.gen.maskedCells.
Next-scale reveals coarse cells first and refines outward; masked reveals a seeded permutation, so both sides of the image are in context at every step. Toggle the order and replot; the target image never changes.
Diffusion heads on an LM trunk
Predict a velocity, not a token id
Discretizing images is one way to fit them into a language model; it is not the only one. Discretizing here means replacing each continuous patch vector with one of a finite set of ids, the route the previous steps took. MAR and Transfusion-style models keep the image latent continuous instead — the latent is the internal vector the trunk works with, a grid of real numbers with no rounding — and attach a small diffusion head to the same trunk. Where a language head emits logits, the raw scores that a softmax turns into probabilities over the text vocabulary, the diffusion head regresses a noise or a velocity, meaning it predicts a vector of numbers by shrinking the distance to a target rather than picking a category. Generation at inference is then not a choice of token ids but a short integration from noise to a clean latent, following the regressed direction step by step. The elegant part is that the trunk itself is untouched: it is still a transformer over a mixed sequence, and the head interprets the last hidden state of an image-latent position as the parameters of a small denoiser — a network that estimates how to clean a noisy vector — rather than as a categorical distribution over a fixed set of token ids. The practical consequence is that the image side never has to commit to a finite vocabulary at all, which sidesteps the collapse problem of the previous step entirely.
Flow matching makes the regression especially clean. Picture sliding a sample along straight rails at constant speed from noise to data. The latent is carried along a straight path from a noise sample $x_0$ to the target $x_1$, the model regresses the constant velocity $x_1 - x_0$ along that path, and sampling is a handful of Euler steps along a line rather than a hundreds-step stochastic denoise. Euler steps are the straightforward way to advance along a direction in small hops, and a stochastic denoise injects fresh randomness at every step; fewer deterministic steps make generation fast and predictable. Because the path is straight and the velocity is constant along it, the model only has to get one vector right rather than a schedule of decreasing noise levels. Classifier-free guidance then nudges the prediction away from the unconditional one — the prediction made with no prompt at all — so the sample follows the prompt more strongly. The plot traces the conditional path, the unconditional path and the guided blend, with the guidance weight on the slider.
Flow paths from Diffusion.flowX with the guided blend of Diffusion.cfg. The conditional and unconditional paths share the flow time axis; guidance extrapolates past the conditional path.
Why the objectives fight
Commitment against smoothness
Understanding is trained by cross-entropy on the next token. Its gradient — the direction in which the weights should move to lower the loss — is largest when the model is wrong about a discrete choice, so it pushes the representation toward tokens that are unambiguous: one right answer, sharply separated from the rest. Generation is trained by a reconstruction or flow objective over a continuous latent, and it is happiest when the latent space is smooth, so that a small perturbation of the noise is a small perturbation of the image rather than a jump to a different object. Those two preferences are not identical. A representation sharp enough to make every next-token prediction confident is often too peaky to sample from without artefacts: a small change in the noise moves the sample across a boundary and the picture changes identity. A representation smooth enough to make sampling stable spreads probability across many nearby possibilities, which is exactly the uncertainty cross-entropy would rather collapse. Take one position's worth of latent: cross-entropy wants it to be unmistakably one thing, while flow matching wants a small nudge to it to produce a slightly different but still coherent picture.
The plot makes the trade a single axis: how discrete the shared representation is, running from smooth and continuous at one end to sharply committed at the other. Move the marker and read the two objective values. They cross in the middle, and no single setting is best for both. In practice a unified model is a negotiated settlement — separate heads, separate losses, a tokenizer chosen to be as smooth as the language side tolerates, and sometimes separate parameter subsets for the two directions. Each of those devices gives one objective somewhere to live without forcing the other to accept it everywhere. The point is not that the tension is intolerable; it is that it is structural, and every unified architecture is a position on this line.
Those devices are exactly the positions that shipped. Chameleon and Emu3 take the pure-autoregressive route: they quantise images into the same kind of discrete tokens as text, concatenate everything into one token stream, and train it with a single next-token loss, so the understanding objective simply dominates and the generation side has to live with the peaks it would rather smooth. Janus-Pro goes the other way and decouples the encoder: the representation the model reads for understanding and the one it generates from are separate, so each direction gets a representation shaped for it, which is the page's own "sometimes separate parameter subsets" line made architectural rather than a promise to train one shared space harder. Show-o keeps one trunk but fuses the two objectives inside it: text is predicted autoregressively while image tokens are produced by masked discrete diffusion — filling in masked positions from the visible ones, the generation order the earlier step showed — so the generation signal is a diffusion objective over discrete codes rather than over a continuous latent, the same one-trunk settlement as Transfusion but reached from the discrete-token side. None of these dissolves the tension: each is a different position on the same discreteness line.
The two objectives against the discreteness of the shared representation. Cross-entropy prefers commitment; the sampling objective prefers smoothness. The marker is the setting selected on the slider.
Further reading
These papers define the tokenizer that made discrete generation work, the variant that removed collapse, and the two generation orders that fixed the conditioning problem, together with the approaches that keep images continuous. Read together they trace the same tension from the codebook outward: first the vocabulary that made image tokens possible, then the attempts to make that vocabulary behave or to do without it, and finally the generation orders and continuous heads that try to reconcile the two directions.
- Aaron van den Oord, Oriol Vinyals and Koray Kavukcuoglu, "Neural Discrete Representation Learning", 2017 — VQ-VAE, the learned codebook and the commitment loss whose failure mode is collapse.
- Fabian Mentzer, David Minnen, Eirikur Agustsson and Michael Tschannen, "Finite Scalar Quantization: VQ-VAE Made Simple", 2023 — FSQ, the fixed per-coordinate grid that cannot collapse because it has no table.
- Keyu Tian, Yi Jiang, Zehuan Yuan and coauthors, "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction", 2024 — VAR, coarse-to-fine scales and parallel prediction within a scale.
- Tianhong Li, Yonglong Tian, He Li and coauthors, "Autoregressive Image Generation without Vector Quantization", 2024 — MAR, masked parallel prediction with a diffusion head and no discrete image vocabulary.
- Chunting Zhou, Lili Yu, Arun Babu and coauthors, "Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model", 2024 — one trunk trained with next-token prediction on text and diffusion on images.
- Chameleon Team, "Chameleon: Mixed-Modal Early-Fusion Foundation Models", 2024 — a token-based trunk over interleaved text and images, and the training instability that discrete image tokens cause.
- Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo and coauthors, "Emu3: Next-Token Prediction is All You Need", 2024 — one token stream and next-token prediction for text, image and video, with no separate diffusion or understanding path.
- Xiaokang Chen, Zhiyu Wu, Xingchao Liu and coauthors, "Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling", 2025 — decoupled encoders for understanding and generation, so each direction reads its own representation.
- Jinheng Xie, Weijia Mao, Zechen Bai and coauthors, "Show-o: One Single Transformer to Unify Multimodal Understanding and Generation", 2024 — autoregressive text and masked discrete diffusion over image tokens fused in one trunk.
Cheat sheet
| Term | Meaning here |
|---|---|
| Unified model | One trunk that both reads images to text and generates images from text |
| VQ-VAE | Autoencoder whose encoder vectors snap to a learned codebook of discrete entries |
| Codebook collapse | A learned codebook concentrating onto a few entries, leaving most of its capacity dead |
| Commitment loss | The term that pulls encoder outputs and codebook entries together, and drives collapse |
| FSQ | Finite scalar quantization: round each latent coordinate on a fixed grid, no codebook |
| Raster order | Row-major autoregression; nearest neighbours in space are far apart in sequence |
| Next-scale (VAR) | Predict whole scales coarse to fine; sequential steps grow logarithmically, not quadratically |
| Masked (MAR) | Reveal positions in a seeded order and predict masked positions in parallel, both sides in context |
| Diffusion head | A small denoiser on the trunk's image-latent positions, regressing noise or velocity |
| Flow matching | Regress the constant velocity along a straight noise-to-sample path; sample with Euler steps |
| Classifier-free guidance | Blend conditional and unconditional predictions to steer the sample toward the prompt |
| Objective conflict | Cross-entropy wants sharp discrete tokens; sampling wants a smooth continuous latent |