Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The three stages

Warm up the adapter, open the language model, then instruct

Stage one is the bridge's turn. The vision tower and the language model are both frozen, and only the bridge is trained: either the projector, a small module that maps each vision vector into the language model's own width, or the Q-Former, a small transformer that reads the patches with a fixed set of learned queries and returns a handful of vectors. The data is usually a few hundred thousand caption pairs, where each pair is one image together with a sentence describing it. The point is not to teach anything new; it is to bring the bridge's random initialisation — the random weight values it starts from before any training at all — into a regime where its gradients are small and sensible before those gradients are allowed to flow into the language model. A gradient is the training signal that says how each of those weights should change, so a large, structured gradient is a violent instruction aimed at a model that was working fine. Skipping this step is one of the most common ways to destroy a good checkpoint, where a checkpoint is a saved set of trained weights. The reason is that a randomly initialised projector produces large, structured errors, and backpropagating them — sending those error signals backward through the network so they update weights — into a frozen-but-now-trainable language model is an efficient way to erase its text ability in a few hundred steps.

Stage two unfreezes the language model, meaning its weights are now allowed to update, and trains it together with the bridge on a much larger mixture — a blend of several datasets rather than a single one. That blend includes captions, visual question answering (questions about an image, paired with short answers), OCR (reading text that appears inside the image), and grounding (answering with a location, such as which region the named object occupies). This is where the model actually learns to use the image, and it is also where forgetting happens. The mechanism is worth naming: because the language model's weights are now moving on data that is mostly images, nothing in the objective is protecting the text knowledge those weights encoded before, so that knowledge can quietly erode. The literature calls this catastrophic forgetting. Stage three is instruction tuning on a smaller, higher-quality set of visual dialogue, where each item is a request paired with the reply it should produce. It is usually run with the vision tower still frozen, and it is what turns a captioning and answering system into an assistant. Some recipes unfreeze the vision tower for the last stage; the toggle below lets you see what that costs.

The schedule: one row per component, one column per stage, filled where the component is trainable and hatched where it is frozen. Stage widths follow the sliders.

💡 By the end of this part you'll be able to describe the three stages and what each trains, say which components are frozen in each, read a data mixture, explain catastrophic forgetting of text ability and why the warm-up exists, say what multimodal DPO adds to the recipe, and say how group-relative reinforcement learning sharpens a model on questions whose answers can be checked. (DPO is a preference-tuning method that learns from pairs of better and worse answers instead of from one correct answer.) These are also the questions worth asking of any recipe you meet, not only this one.
2

What is frozen when

Three components, three schedules

The freezing schedule — the per-stage list of which components are allowed to learn and which are held fixed — is the schedule's whole content, and the pattern in this demo is the LLaVA shape, named after the model that made it standard. The projector is trainable in every stage. The language model is frozen in stage one and trainable from stage two onward. The vision tower is frozen throughout unless the toggle is on, in which case it joins from stage two. Because the language model is a billion-parameter object — a parameter being one learned number in the network, so a billion of them is roughly the size of the whole text model — and the projector is a few tens of millions, the jump in trainable parameters between stage one and stage two is a factor of hundreds. Put a number on it and the reason for the separate first step is plain: warming up the bridge alone moves tens of millions of weights, while the opening step of stage two moves billions. That is exactly why stage one must exist as a separate, gentle step rather than as the first iteration of stage two.

There is a second reason the pattern matters. The language model carries text ability that was expensive to obtain: it reasons, follows instructions and writes fluently, and those skills were paid for with a very large text corpus. There is no stage in this recipe whose data is mostly text, so the training signal is dominated by images. That matters because any step that moves the language model's weights is a step that can move them away from fluency, reasoning and instruction following, and the data mixture — the deliberate blend of text and image examples — is the only thing holding that line. Nothing in the multimodal objective marks the text weights as off-limits, so they are free to move until something else holds them in place. The bars below price each stage; the curve in the next-but-one step is what happens to the text side of the ledger, the running account of what the original model could still do.

Trainable parameters by stage, in billions, under the freeze pattern above. The projector is a rounding error next to the language model, which is the point of warming it up alone.

Rough sizes: projector 21 M, language model 7.0 B, vision tower 0.30 B. The numbers scale with the choice in the previous part.

⚠ The unfreeze decision is the expensive one. Training the vision tower multiplies the vision-side compute and adds a second source of catastrophic forgetting, because the encoder may lose the general visual features that a contrastive pretraining run paid for — contrastive pretraining being the earlier stage that taught the encoder to match images with their captions, which is where its broad visual knowledge comes from. Most recipes leave it frozen and spend the budget on data instead.
3

The data mixture

Captions teach alignment, VQA teaches use, instructions teach behaviour

The data changes with the stage because the objective — the quantity the model is being trained to improve — does. Stage one wants pairs that are easy and unambiguous, so it uses caption data, where the only thing to learn is which image vector corresponds to which caption. Stage two needs the model to extract information the captions never mentioned, such as the text on a sign or where a small object sits, so it mixes in visual question answering, OCR and grounding sets, and it deliberately keeps some text-only data to slow the drift in language ability. A caption is a weak teaching signal for a specific question, because the same photo can be described in many correct ways; a question-answer item instead asks one thing and expects a short, definite reply, which is why stage two needs it. Stage three is small and curated, meaning every example was chosen or written on purpose rather than scraped in bulk: the answers are long and conversational, and the point is the format of the interaction rather than the volume of the signal.

StageTypical mixtureScaleWhat it teaches
1 · warm-upCaption pairs; CC3M, LAION, alt text~0.5 M pairsProject the vision features into the language model's space
2 · pretrainingCaptions + VQA + OCR + grounding + a little text-only~1 M+ samplesGround answers in the image, and stop the language model drifting
3 · instructionShort visual dialogues, detailed descriptions, refusals~0.5 M turnsFollow instructions, answer at the right length, admit uncertainty
4 · preferencePairs of responses to the same image, one preferred~10 K–100 K pairsPrefer the careful answer over the confident wrong one
💡 The mixture is where the deployment shows up. A model trained mostly on short captions is a strong captioner and a weak assistant; a model trained mostly on long instruction answers becomes verbose and starts describing things that are not there. When a failure looks like a behaviour, look at the mixture before looking at the architecture. That order matters because the mixture is cheap to change and the architecture is not, so it is usually the first lever worth pulling.
4

What breaks

The trade the schedule is trying to keep

Turning up the language model's learning rate — the step size training takes when it updates weights, so a larger value moves them further per example — or leaving it unfrozen for too long without text data, produces a characteristic curve: visual question answering improves while text benchmarks fall. A benchmark is a fixed test set used to score a model, so a falling number means measured ability, not one unlucky answer. The model is being optimised for the multimodal objective, and the text ability that the objective does not directly reward is free to erode. The failure is quiet, because the model still produces fluent sentences; it has lost some of the reasoning and world knowledge that made those sentences correct. Picture two runs at the same budget: one spends every batch on images and scores higher on visual question answering, while one keeps a quarter of its batches as plain text and scores a little lower on images but holds its reasoning. Mixing in text-only data, lowering the learning rate, or freezing for more of training all trade multimodal accuracy for retention, and no recipe escapes the trade entirely.

Two other failures are worth naming because they look like bugs and are not. The first is that OCR and small-text reading remain weak long after captioning looks solved. The reason is resolution: characters are small and fine, and the default pipeline — the fixed chain of resizing and patching that turns a picture into tokens — has already thrown that fine detail away before the language model sees anything, which is the previous part's compression in action. The second is that grounding, the ability to point at the object it just named, lags behind naming it. Naming only needs the right features to exist somewhere in the image token set, while grounding needs them in the right place: a model can tell that a dog is present from a blurred patch and still not be able to say which patch it occupies.

A sketch of the trade: VQA accuracy rises and text benchmark accuracy falls as the language model is tuned harder on multimodal data. The marker is the current setting.

The curves are illustrative, not measured: the shape is the claim, plus the fact that the two lines cross rather than rise together.

⚠ Text-only data is the standard antidote. Keeping a fraction of pure-language batches in the mixture is the cheapest way to hold the text side of the ledger, and it costs image throughput rather than model capacity — throughput being how many examples the run gets through per unit of time, so a text batch is spent on language instead of images and the image side learns a little slower. Recipes that drop it for speed report exactly the forgetting curve above.
5

Preference tuning on pairs

Learning from better and worse, not from right and wrong

Supervised tuning — also called supervised fine-tuning, or SFT — needs a reference answer for every prompt, and for a multimodal question there is often no single right answer. Ask ten people to describe a photograph and you get ten different sentences; none is wrong, so treating one of them as the target teaches the model to imitate that style rather than to prefer a better answer. Preference tuning takes a different route: it asks for two responses, marks one preferred, and trains the model to raise the likelihood of the chosen one relative to the rejected one. No reference answer is needed, only a judgement about which of the two is better. In multimodal DPO — direct preference optimisation, the objective that adapts this idea to image-conditioned pairs — the pair shares the same image. Because the two answers look at the same pixels, the difference between them is usually carefulness rather than perception, so what is learned is to stop over-claiming. A reference model — a frozen copy of the model from before this stage — anchors the update, keeping the trained model from wandering too far, and the coefficient $\beta$ controls how far the policy, meaning the model being trained, is allowed to drift from it. Small $\beta$ learns preferences fast but drifts far and reward-hacks, meaning it exploits the preference signal rather than genuinely getting better; large $\beta$ stays close and barely moves.

The plot below shows the two failure modes as functions of $\beta$: held-out preference accuracy — scores on pairs the training run never saw, so the number measures generalisation rather than memorisation — which rises as the model is allowed to move, and the rate at which it drifts into confidently wrong answers. The sweet spot is the region where accuracy has risen but hacking has not yet climbed — and it is narrow enough that $\beta$ is treated as a real hyperparameter, a setting you tune by experiment and defend with evidence, rather than as a formality.

Held-out preference accuracy and the rate of confident hallucination against the DPO coefficient β. Both are fractions, so they share one axis.

Small β moves far from the reference model; large β barely moves. The hallucination curve is what makes the small-β end dangerous rather than merely different.

6

RL and reasoning post-training

The group is its own baseline

Preference tuning taught the model which of two answers a person would rather have. A newer stage asks a sharper question: among all the answers the model could produce, which ones are actually correct? That is a reinforcement-learning problem, and reinforcement learning here means training by trial and reward rather than by imitation: the model generates its own attempts, something scores them, and the update makes the better attempts more likely and the worse ones less likely. The immediate difficulty is that a reward is only a number, and a number means nothing on its own. A score of 0.6 is good or bad depending on what the model usually gets on that prompt, so every method in this family needs a baseline — a reference point for what an ordinary attempt scores — because without one the update cannot tell a genuine improvement from a lucky draw. Proximal policy optimisation, or PPO, the method that dominated reinforcement learning for language models, learns that baseline with a second network called a critic, trained to predict the reward the policy is about to receive. The critic roughly doubles the memory, needs its own tuning, and in a vision-language model it must also learn to read the image, because the reward it predicts depends on what was in the picture. That is a great deal of machinery to bolt onto the last stage of a recipe.

Group relative policy optimisation, or GRPO, removes the critic by noticing where a baseline can be had for free. For a single prompt, sample not one response but a whole group of them — typically four to sixteen attempts at the same question — and score every attempt in the group with the reward function. The group's own average score is then the baseline. Each attempt's advantage, the quantity the update actually consumes, is its reward minus the group mean, divided by the group's spread, so that a response better than its siblings gets a positive advantage and its tokens are made more likely, while a response worse than its siblings gets a negative advantage and is suppressed. No critic network, no value head, no second model to train: the group is the baseline, and the reward scale cancels out of the ratio entirely. That one substitution is why GRPO spread through the open-model releases of 2025, and it was the method behind DeepSeek-R1's reasoning behaviour. The reward has to be computable without a person in the loop, because the group must be scored automatically and repeatedly. For a text model the checkable rewards are mathematics and code, where the final answer can be verified by a parser or a test run; for a vision-language model they are precisely the tasks the earlier parts of this volume set up: a chart question with a known answer, a document question that must quote the right field, a bounding box scored by how much it overlaps the true one, or a visual arithmetic problem whose final number can be read and compared. That recipe — reinforcement learning against a rule that can be checked, rather than against a human preference — is what produces the “thinking” vision-language models, which write out a chain of intermediate reasoning before committing to an answer. On the text side the proof was DeepSeek-R1; among open multimodal models, GLM-4.1V-Thinking and the reasoning-tuned Qwen3-VL variants are the same move applied to images.

The plot below is the mechanism, not the reward. It shows one prompt's group of sampled attempts on a reward axis, the group mean as the dashed line they are measured against, and each attempt's deviation from that line as a bar. The slider sets how many attempts the group contains, which is the method's main tuning knob and its main cost. With only two attempts the baseline is the average of two samples, so one lucky draw receives an enormous advantage and the update chases it; as the group grows the mean becomes a stabler estimate of what this prompt is worth to the model, the advantages shrink toward the group's genuine spread, and the signal is cleaner. The price is that every attempt in the group is a full generation, so a group of sixteen costs eight times the sampling compute of a group of two for the same number of prompts. Move the slider and watch the readout's advantages settle: at two attempts they are pinned near plus and minus one, and by sixteen the outliers separate from the pack.

One prompt's sampled group on a reward axis, the group mean as the baseline, and each attempt's signed distance from that mean. The advantages in the readout come from Multimodal.rl.groupAdvantage.

Advantage = (reward minus group mean) / group spread. Small groups give a noisy baseline; large groups give a stable one and cost more samples per prompt.

💡 Where this stage sits. Supervised tuning teaches the format of an answer, preference tuning teaches which of two answers is more careful, and group-relative reinforcement learning teaches correctness on questions that can be checked. They stack rather than replace one another, and the open recipes that produce a reasoning model run the earlier stages first, because a model that cannot follow the answer format will not reliably emit something a verifier can read. The cost of the stage is verbosity: a chain of intermediate reasoning is many more tokens and a longer wait, a deliberate trade of latency for accuracy on the hard cases.
⚠ An uninformative reward is worse than none. Group-relative advantages depend entirely on the reward being able to rank attempts within the group. If every attempt scores the same — because the verifier is too coarse, or the question too easy, or the answers too similar — then the group's spread is zero and every advantage is zero, so no gradient flows at all. A reward that is merely noisy is worse still, because it ranks attempts at random and the update reinforces whichever attempt happened to be scored highest. Getting the reward right, not the optimiser, is the hard part of this stage.

Further reading

The staging recipe, the instruction-tuning data, and the methods that followed it. Read them in this order and the arc runs from the three-stage schedule, through the mixture and unfreezing decisions that refine it, to the objectives that teach a trained model to prefer a better answer, and finally to the reinforcement-learning stage that teaches it to check one.

Cheat sheet

TermMeaning here
Warm-up stageTrain the bridge alone against a frozen language model on caption data
Joint pretrainingUnfreeze the language model, train it with the bridge on a mixed multimodal corpus
Instruction tuningSmall, curated visual dialogues that turn a captioner into an assistant
Catastrophic forgettingText ability degrades as the language model is tuned on image-dominated data
Freezing scheduleWhich components receive gradients in which stage; the schedule's real content
Data mixtureThe per-stage blend of captions, VQA, OCR, grounding and text-only data
Grounding lagNaming an object is easier than pointing at it, so grounding trains last and lags
Preference tuningTrain on chosen/rejected pairs rather than on a single reference answer
DPODirect preference objective; $\beta$ sets how far the policy may drift from the reference
RL post-trainingSample attempts, score them, and update toward the better ones; no reference answer
GRPOGroup-relative policy optimisation: advantage = (reward minus group mean) / group spread, no critic
Verifiable rewardA score a rule can compute — a checked answer, a matched box — instead of a human preference
Thinking VLMA model that writes intermediate reasoning before answering, bought with RL and paid for in latency
7

Check your understanding

0/5 answered