Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

One model for both

The same trunk, two directions

A vision-language model maps pixels to words through a bridge — the small module that translates one tower's vectors into the other's format — and a language model. A unified model keeps one trunk and asks it to do the mapping both ways: image tokens in, text tokens out for understanding, which here means reading a picture and answering questions about it; text plus image tokens in, more image tokens out for generation, which means producing a new picture rather than describing an existing one. Picture a single model that can caption a photograph and then, handed the caption "a red mug on a table", draw a fresh picture of one. The attraction is not that the architecture is elegant — it is that understanding and generation are complementary training signals for the same representation. They exercise different skills on the same underlying content, so together they push the trunk to learn more than either could alone. A model that must draw an object has to learn how its parts fit together and how it looks from different angles; a model that only answers questions about it does not have to learn any of that. Think of the two tasks as two ways of studying the same object, describing it and reconstructing it from memory, where the second exposes structure the first can skate past.

What makes this hard is that the two directions want different things from the token stream, the sequence of tokens the model reads and writes. Understanding wants a token whose distribution can be made sharp, because cross-entropy rewards confidence: it gives a low loss only when the model puts most of its probability on the one right answer. Generation wants a token that is easy to sample from and whose neighbours are visually plausible, because here there is no single right answer and a single wrong jump is a visible artefact — a smudge, a broken edge, an object that suddenly changes identity. Those two preferences do not point the same way. When the tokens are discrete and learned, so that each image token is one of a finite set of categories the model learned, the two pull on the same quantizer, the component that decides which category each piece of image becomes; when the trunk predicts continuous latents with a diffusion head, it stops predicting a category at all. A quantizer tuned for confidence would rather make the gap between neighbouring codes enormous, so the right id is never in doubt; a generator would rather the codes it samples blend smoothly into nearby ones, so a small mistake is not visible. The rest of the part is about that tension, starting with the tokenizer at the front.

💡 By the end of this part you'll be able to explain why a codebook can collapse while FSQ cannot, describe raster, next-scale and masked generation orders and what each conditions on, say what a diffusion head predicts instead of a token id, and name the reason the understanding and generation objectives fight. Each of those is a piece of one story: how a single vocabulary is shared between two jobs that want different things from it.
2

The tokenizer crux

A learned codebook can collapse; a fixed grid cannot

An autoencoder turns an image into a grid of continuous vectors — an encoder compresses the picture into those vectors, and a decoder reconstructs the picture back out of them — but a language-model trunk wants discrete ids, whole integers from a countable vocabulary. The classic bridge between those two worlds is a vector-quantized variational autoencoder. Think of tokenizing an image as cutting a photo into a contact sheet of equal tiles and reading them left to right, top to bottom, like words. Vector quantisation is the rule that gives each tile one word: the encoder vector for a tile is snapped to its nearest entry in a learned codebook, a fixed-size table of vectors that acts as the vocabulary, and the decoder reconstructs from those entries rather than from the original vectors. Training adds a commitment loss that pulls encoder outputs toward codebook entries and entries toward encoder outputs, so the two meet in the middle. Concretely, two tiles of sky that look similar have encoder vectors that sit close together, and both may snap to the same entry; the table is what turns an unbounded set of vectors into a countable set of words. The failure mode is exactly what that coupling invites. Early in training a few entries are slightly closer, they win more vectors, they get pulled closer still, and the rest of the codebook becomes dead weight. The model still reconstructs, using a handful of entries, and the discrete bottleneck — the narrow channel of discrete choices every tile must pass through — has quietly become a coarse quantizer with most of its capacity unused. That is codebook collapse, and it degrades generation most: sampling has fewer distinct tokens to compose with, so the model has a smaller vocabulary for building an image than it was supposed to.

FSQ — finite scalar quantization — removes the thing that collapses. Instead of a learned table, each coordinate of the latent is rounded against a fixed set of levels. If a tile's latent vector has three coordinates and each is rounded to one of four levels, the code is just the tuple of those three per-coordinate indices, and there are 4 × 4 × 4 = 64 possible codes. There is no nearest-neighbour search, no commitment loss and no table to starve: every level is reachable by construction, and the implicit vocabulary size is the product of the level counts. The codebook still exists as a mathematical set of codes, but it is implied by the grid rather than learned, so nothing in training can pull usage toward a few favoured entries. The trade is that a fixed grid cannot adapt to the data the way a learned table can: it is safe, but it never discovers entries that the images would have rewarded. The bars below show what happens to a codebook's usage as an auxiliary pressure is turned up. The learned codebook concentrates onto a few entries; the fixed grid stays close to uniform, because nothing in its objective can reward concentration.

Codebook entry usage for a learned VQ codebook and for a fixed FSQ grid, normalised by the busiest entry. The collapse slider raises the pressure that drives a VQ table to concentrate; the FSQ bars do not move.

⚠ Collapse is not a bug you can tune away. Rotation tricks, code reinitialisation and commitment schedules all slow it, but the mechanism — a learned table coupled to its own usage — remains. Removing the table removes the mechanism, which is why FSQ and its relatives are attractive for a unified model that needs its full vocabulary alive. Each dead entry is capacity the generator paid for and cannot use: fewer live tokens means fewer distinct tile patterns the model can ever produce, so the pictures it makes are coarser than its parameter count suggests.
3

Three generation orders

What the model conditions on while it decides

A raster (row-major) order reads image tokens like a sentence: top-left to bottom-right, one token at a time, each conditioned on everything before it. Conditioned on means the model's prediction for the current token is allowed to look at every token already produced. It is the order a plain autoregressive transformer already understands — autoregressive is the standard recipe of predicting the next token from the ones before it, one step at a time — and it is the weakest for images, because a token's nearest neighbours in space are far away in the generation order. The token below a pixel is generated hundreds of steps after the token above it, so the model has to remember a lattice in a one-dimensional stream. By the time it reaches the bottom row, the top row is far behind in the sequence and the spatial structure it needs is no longer nearby. Raster is nevertheless the order a text model can adopt with no architectural change, which is exactly why it was tried first: the sequence is already one-dimensional, and next-token prediction needs no modification at all.

VAR — visual autoregressive modelling — changes the axis rather than the mechanism. It generates the image as a sequence of scales, meaning resolutions: a single coarse token for the whole picture first, then a $2 \times 2$ grid, then $4 \times 4$, each scale conditioned on all coarser scales. Picture an artist blocking in a painting: first the rough shape and colour of the whole scene, then progressively finer strokes, each consistent with the block-in underneath. Because the next scale is predicted in parallel, with all the tokens of a scale produced at once rather than one at a time, the number of sequential steps is logarithmic in resolution instead of quadratic, and every token at a scale sees the global structure before it commits to detail. In other words, doubling the resolution adds a step rather than quadrupling the work. For the four scales 1×1 through 8×8, that is four sequential passes rather than eighty-five token steps. MAR — masked autoregressive — takes the opposite route to the same goal. It reveals positions in a seeded order, a pseudo-random sequence fixed by a seed so it can be reproduced, and at each step predicts a whole set of masked positions at once, conditioned on all revealed ones. Because the revealed tokens can sit on every side of a masked one, the model can use context on both sides of a token, not just the tokens before it. The buttons switch the reveal order on a small grid and animate it; the picture is the same, only the order in which it is committed changes.

A $32 \times 32$ test image drawn with GenMedia.raster and revealed as an $8 \times 8$ grid of cells. The order comes from Multimodal.gen.rasterCells, Multimodal.gen.nextScaleCells and Multimodal.gen.maskedCells.

Next-scale reveals coarse cells first and refines outward; masked reveals a seeded permutation, so both sides of the image are in context at every step. Toggle the order and replot; the target image never changes.

💡 Order is a conditioning choice, not a cosmetic one. Reducing the number of sequential steps cuts inference latency; changing what is in context changes what the model can get right. Next-scale and masked orders are both attempts to give each decision more of the image than its predecessors in raster order. Latency here is the wall-clock time to produce a whole image, and the difference between one pass over a few scales and thousands of sequential token steps is the difference between a fraction of a second and many seconds.
5

Why the objectives fight

Commitment against smoothness

Understanding is trained by cross-entropy on the next token. Its gradient — the direction in which the weights should move to lower the loss — is largest when the model is wrong about a discrete choice, so it pushes the representation toward tokens that are unambiguous: one right answer, sharply separated from the rest. Generation is trained by a reconstruction or flow objective over a continuous latent, and it is happiest when the latent space is smooth, so that a small perturbation of the noise is a small perturbation of the image rather than a jump to a different object. Those two preferences are not identical. A representation sharp enough to make every next-token prediction confident is often too peaky to sample from without artefacts: a small change in the noise moves the sample across a boundary and the picture changes identity. A representation smooth enough to make sampling stable spreads probability across many nearby possibilities, which is exactly the uncertainty cross-entropy would rather collapse. Take one position's worth of latent: cross-entropy wants it to be unmistakably one thing, while flow matching wants a small nudge to it to produce a slightly different but still coherent picture.

The plot makes the trade a single axis: how discrete the shared representation is, running from smooth and continuous at one end to sharply committed at the other. Move the marker and read the two objective values. They cross in the middle, and no single setting is best for both. In practice a unified model is a negotiated settlement — separate heads, separate losses, a tokenizer chosen to be as smooth as the language side tolerates, and sometimes separate parameter subsets for the two directions. Each of those devices gives one objective somewhere to live without forcing the other to accept it everywhere. The point is not that the tension is intolerable; it is that it is structural, and every unified architecture is a position on this line.

Those devices are exactly the positions that shipped. Chameleon and Emu3 take the pure-autoregressive route: they quantise images into the same kind of discrete tokens as text, concatenate everything into one token stream, and train it with a single next-token loss, so the understanding objective simply dominates and the generation side has to live with the peaks it would rather smooth. Janus-Pro goes the other way and decouples the encoder: the representation the model reads for understanding and the one it generates from are separate, so each direction gets a representation shaped for it, which is the page's own "sometimes separate parameter subsets" line made architectural rather than a promise to train one shared space harder. Show-o keeps one trunk but fuses the two objectives inside it: text is predicted autoregressively while image tokens are produced by masked discrete diffusion — filling in masked positions from the visible ones, the generation order the earlier step showed — so the generation signal is a diffusion objective over discrete codes rather than over a continuous latent, the same one-trunk settlement as Transfusion but reached from the discrete-token side. None of these dissolves the tension: each is a different position on the same discreteness line.

The two objectives against the discreteness of the shared representation. Cross-entropy prefers commitment; the sampling objective prefers smoothness. The marker is the setting selected on the slider.

💡 The settlement is visible in the architecture. Separate output heads, auxiliary reconstruction losses, tokenizers with a large but not-too-peaky vocabulary, and mixture-of-experts trunks that let some capacity specialise for generation are all ways of buying room on both axes instead of picking a point. Read any unified design and you can usually say which objective each part was added to placate.

Further reading

These papers define the tokenizer that made discrete generation work, the variant that removed collapse, and the two generation orders that fixed the conditioning problem, together with the approaches that keep images continuous. Read together they trace the same tension from the codebook outward: first the vocabulary that made image tokens possible, then the attempts to make that vocabulary behave or to do without it, and finally the generation orders and continuous heads that try to reconcile the two directions.

Cheat sheet

TermMeaning here
Unified modelOne trunk that both reads images to text and generates images from text
VQ-VAEAutoencoder whose encoder vectors snap to a learned codebook of discrete entries
Codebook collapseA learned codebook concentrating onto a few entries, leaving most of its capacity dead
Commitment lossThe term that pulls encoder outputs and codebook entries together, and drives collapse
FSQFinite scalar quantization: round each latent coordinate on a fixed grid, no codebook
Raster orderRow-major autoregression; nearest neighbours in space are far apart in sequence
Next-scale (VAR)Predict whole scales coarse to fine; sequential steps grow logarithmically, not quadratically
Masked (MAR)Reveal positions in a seeded order and predict masked positions in parallel, both sides in context
Diffusion headA small denoiser on the trunk's image-latent positions, regressing noise or velocity
Flow matchingRegress the constant velocity along a straight noise-to-sample path; sample with Euler steps
Classifier-free guidanceBlend conditional and unconditional predictions to steer the sample toward the prompt
Objective conflictCross-entropy wants sharp discrete tokens; sampling wants a smooth continuous latent
7

Check your understanding

0/5 answered