Training a vision-language model
Nobody trains a vision-language model from scratch. A vision-language model is one that reads an image and writes text about it, and the surprising part is that neither of its two large halves is built here. The vision tower — the half that turns pixels into vectors — arrives pretrained, meaning someone else already trained it on a large pile of images, and the language model arrives pretrained too, meaning it already knows how to read and write text. The whole job is to teach those two finished parts to work together without wrecking either one. A recipe is the plan that does this, and it is run in stages: separate training phases, one after another, each with its own settings. First the bridge — the small learned module that translates vision vectors into vectors the language model accepts — is warmed up against a frozen language model, where frozen means its weights are held fixed and receive no updates. Then the language model is opened up, so its weights can finally move, and trained with the bridge on a large mixed corpus, a big collection of examples that blends several kinds of task. Finally a smaller instruction set turns the result into something that follows directions. Each stage has its own freezing pattern — its own answer to which parts are allowed to learn — its own data, and its own characteristic failure. This part walks the schedule, prices its stages, and shows what happens when it is run too hard.
The three stages
Warm up the adapter, open the language model, then instruct
Stage one is the bridge's turn. The vision tower and the language model are both frozen, and only the bridge is trained: either the projector, a small module that maps each vision vector into the language model's own width, or the Q-Former, a small transformer that reads the patches with a fixed set of learned queries and returns a handful of vectors. The data is usually a few hundred thousand caption pairs, where each pair is one image together with a sentence describing it. The point is not to teach anything new; it is to bring the bridge's random initialisation — the random weight values it starts from before any training at all — into a regime where its gradients are small and sensible before those gradients are allowed to flow into the language model. A gradient is the training signal that says how each of those weights should change, so a large, structured gradient is a violent instruction aimed at a model that was working fine. Skipping this step is one of the most common ways to destroy a good checkpoint, where a checkpoint is a saved set of trained weights. The reason is that a randomly initialised projector produces large, structured errors, and backpropagating them — sending those error signals backward through the network so they update weights — into a frozen-but-now-trainable language model is an efficient way to erase its text ability in a few hundred steps.
Stage two unfreezes the language model, meaning its weights are now allowed to update, and trains it together with the bridge on a much larger mixture — a blend of several datasets rather than a single one. That blend includes captions, visual question answering (questions about an image, paired with short answers), OCR (reading text that appears inside the image), and grounding (answering with a location, such as which region the named object occupies). This is where the model actually learns to use the image, and it is also where forgetting happens. The mechanism is worth naming: because the language model's weights are now moving on data that is mostly images, nothing in the objective is protecting the text knowledge those weights encoded before, so that knowledge can quietly erode. The literature calls this catastrophic forgetting. Stage three is instruction tuning on a smaller, higher-quality set of visual dialogue, where each item is a request paired with the reply it should produce. It is usually run with the vision tower still frozen, and it is what turns a captioning and answering system into an assistant. Some recipes unfreeze the vision tower for the last stage; the toggle below lets you see what that costs.
The schedule: one row per component, one column per stage, filled where the component is trainable and hatched where it is frozen. Stage widths follow the sliders.
What is frozen when
Three components, three schedules
The freezing schedule — the per-stage list of which components are allowed to learn and which are held fixed — is the schedule's whole content, and the pattern in this demo is the LLaVA shape, named after the model that made it standard. The projector is trainable in every stage. The language model is frozen in stage one and trainable from stage two onward. The vision tower is frozen throughout unless the toggle is on, in which case it joins from stage two. Because the language model is a billion-parameter object — a parameter being one learned number in the network, so a billion of them is roughly the size of the whole text model — and the projector is a few tens of millions, the jump in trainable parameters between stage one and stage two is a factor of hundreds. Put a number on it and the reason for the separate first step is plain: warming up the bridge alone moves tens of millions of weights, while the opening step of stage two moves billions. That is exactly why stage one must exist as a separate, gentle step rather than as the first iteration of stage two.
There is a second reason the pattern matters. The language model carries text ability that was expensive to obtain: it reasons, follows instructions and writes fluently, and those skills were paid for with a very large text corpus. There is no stage in this recipe whose data is mostly text, so the training signal is dominated by images. That matters because any step that moves the language model's weights is a step that can move them away from fluency, reasoning and instruction following, and the data mixture — the deliberate blend of text and image examples — is the only thing holding that line. Nothing in the multimodal objective marks the text weights as off-limits, so they are free to move until something else holds them in place. The bars below price each stage; the curve in the next-but-one step is what happens to the text side of the ledger, the running account of what the original model could still do.
Trainable parameters by stage, in billions, under the freeze pattern above. The projector is a rounding error next to the language model, which is the point of warming it up alone.
Rough sizes: projector 21 M, language model 7.0 B, vision tower 0.30 B. The numbers scale with the choice in the previous part.
The data mixture
Captions teach alignment, VQA teaches use, instructions teach behaviour
The data changes with the stage because the objective — the quantity the model is being trained to improve — does. Stage one wants pairs that are easy and unambiguous, so it uses caption data, where the only thing to learn is which image vector corresponds to which caption. Stage two needs the model to extract information the captions never mentioned, such as the text on a sign or where a small object sits, so it mixes in visual question answering, OCR and grounding sets, and it deliberately keeps some text-only data to slow the drift in language ability. A caption is a weak teaching signal for a specific question, because the same photo can be described in many correct ways; a question-answer item instead asks one thing and expects a short, definite reply, which is why stage two needs it. Stage three is small and curated, meaning every example was chosen or written on purpose rather than scraped in bulk: the answers are long and conversational, and the point is the format of the interaction rather than the volume of the signal.
| Stage | Typical mixture | Scale | What it teaches |
|---|---|---|---|
| 1 · warm-up | Caption pairs; CC3M, LAION, alt text | ~0.5 M pairs | Project the vision features into the language model's space |
| 2 · pretraining | Captions + VQA + OCR + grounding + a little text-only | ~1 M+ samples | Ground answers in the image, and stop the language model drifting |
| 3 · instruction | Short visual dialogues, detailed descriptions, refusals | ~0.5 M turns | Follow instructions, answer at the right length, admit uncertainty |
| 4 · preference | Pairs of responses to the same image, one preferred | ~10 K–100 K pairs | Prefer the careful answer over the confident wrong one |
What breaks
The trade the schedule is trying to keep
Turning up the language model's learning rate — the step size training takes when it updates weights, so a larger value moves them further per example — or leaving it unfrozen for too long without text data, produces a characteristic curve: visual question answering improves while text benchmarks fall. A benchmark is a fixed test set used to score a model, so a falling number means measured ability, not one unlucky answer. The model is being optimised for the multimodal objective, and the text ability that the objective does not directly reward is free to erode. The failure is quiet, because the model still produces fluent sentences; it has lost some of the reasoning and world knowledge that made those sentences correct. Picture two runs at the same budget: one spends every batch on images and scores higher on visual question answering, while one keeps a quarter of its batches as plain text and scores a little lower on images but holds its reasoning. Mixing in text-only data, lowering the learning rate, or freezing for more of training all trade multimodal accuracy for retention, and no recipe escapes the trade entirely.
Two other failures are worth naming because they look like bugs and are not. The first is that OCR and small-text reading remain weak long after captioning looks solved. The reason is resolution: characters are small and fine, and the default pipeline — the fixed chain of resizing and patching that turns a picture into tokens — has already thrown that fine detail away before the language model sees anything, which is the previous part's compression in action. The second is that grounding, the ability to point at the object it just named, lags behind naming it. Naming only needs the right features to exist somewhere in the image token set, while grounding needs them in the right place: a model can tell that a dog is present from a blurred patch and still not be able to say which patch it occupies.
A sketch of the trade: VQA accuracy rises and text benchmark accuracy falls as the language model is tuned harder on multimodal data. The marker is the current setting.
The curves are illustrative, not measured: the shape is the claim, plus the fact that the two lines cross rather than rise together.
Preference tuning on pairs
Learning from better and worse, not from right and wrong
Supervised tuning — also called supervised fine-tuning, or SFT — needs a reference answer for every prompt, and for a multimodal question there is often no single right answer. Ask ten people to describe a photograph and you get ten different sentences; none is wrong, so treating one of them as the target teaches the model to imitate that style rather than to prefer a better answer. Preference tuning takes a different route: it asks for two responses, marks one preferred, and trains the model to raise the likelihood of the chosen one relative to the rejected one. No reference answer is needed, only a judgement about which of the two is better. In multimodal DPO — direct preference optimisation, the objective that adapts this idea to image-conditioned pairs — the pair shares the same image. Because the two answers look at the same pixels, the difference between them is usually carefulness rather than perception, so what is learned is to stop over-claiming. A reference model — a frozen copy of the model from before this stage — anchors the update, keeping the trained model from wandering too far, and the coefficient $\beta$ controls how far the policy, meaning the model being trained, is allowed to drift from it. Small $\beta$ learns preferences fast but drifts far and reward-hacks, meaning it exploits the preference signal rather than genuinely getting better; large $\beta$ stays close and barely moves.
The plot below shows the two failure modes as functions of $\beta$: held-out preference accuracy — scores on pairs the training run never saw, so the number measures generalisation rather than memorisation — which rises as the model is allowed to move, and the rate at which it drifts into confidently wrong answers. The sweet spot is the region where accuracy has risen but hacking has not yet climbed — and it is narrow enough that $\beta$ is treated as a real hyperparameter, a setting you tune by experiment and defend with evidence, rather than as a formality.
Held-out preference accuracy and the rate of confident hallucination against the DPO coefficient β. Both are fractions, so they share one axis.
Small β moves far from the reference model; large β barely moves. The hallucination curve is what makes the small-β end dangerous rather than merely different.
RL and reasoning post-training
The group is its own baseline
Preference tuning taught the model which of two answers a person would rather have. A newer stage asks a sharper question: among all the answers the model could produce, which ones are actually correct? That is a reinforcement-learning problem, and reinforcement learning here means training by trial and reward rather than by imitation: the model generates its own attempts, something scores them, and the update makes the better attempts more likely and the worse ones less likely. The immediate difficulty is that a reward is only a number, and a number means nothing on its own. A score of 0.6 is good or bad depending on what the model usually gets on that prompt, so every method in this family needs a baseline — a reference point for what an ordinary attempt scores — because without one the update cannot tell a genuine improvement from a lucky draw. Proximal policy optimisation, or PPO, the method that dominated reinforcement learning for language models, learns that baseline with a second network called a critic, trained to predict the reward the policy is about to receive. The critic roughly doubles the memory, needs its own tuning, and in a vision-language model it must also learn to read the image, because the reward it predicts depends on what was in the picture. That is a great deal of machinery to bolt onto the last stage of a recipe.
Group relative policy optimisation, or GRPO, removes the critic by noticing where a baseline can be had for free. For a single prompt, sample not one response but a whole group of them — typically four to sixteen attempts at the same question — and score every attempt in the group with the reward function. The group's own average score is then the baseline. Each attempt's advantage, the quantity the update actually consumes, is its reward minus the group mean, divided by the group's spread, so that a response better than its siblings gets a positive advantage and its tokens are made more likely, while a response worse than its siblings gets a negative advantage and is suppressed. No critic network, no value head, no second model to train: the group is the baseline, and the reward scale cancels out of the ratio entirely. That one substitution is why GRPO spread through the open-model releases of 2025, and it was the method behind DeepSeek-R1's reasoning behaviour. The reward has to be computable without a person in the loop, because the group must be scored automatically and repeatedly. For a text model the checkable rewards are mathematics and code, where the final answer can be verified by a parser or a test run; for a vision-language model they are precisely the tasks the earlier parts of this volume set up: a chart question with a known answer, a document question that must quote the right field, a bounding box scored by how much it overlaps the true one, or a visual arithmetic problem whose final number can be read and compared. That recipe — reinforcement learning against a rule that can be checked, rather than against a human preference — is what produces the “thinking” vision-language models, which write out a chain of intermediate reasoning before committing to an answer. On the text side the proof was DeepSeek-R1; among open multimodal models, GLM-4.1V-Thinking and the reasoning-tuned Qwen3-VL variants are the same move applied to images.
The plot below is the mechanism, not the reward. It shows one prompt's group of sampled attempts on a reward axis, the group mean as the dashed line they are measured against, and each attempt's deviation from that line as a bar. The slider sets how many attempts the group contains, which is the method's main tuning knob and its main cost. With only two attempts the baseline is the average of two samples, so one lucky draw receives an enormous advantage and the update chases it; as the group grows the mean becomes a stabler estimate of what this prompt is worth to the model, the advantages shrink toward the group's genuine spread, and the signal is cleaner. The price is that every attempt in the group is a full generation, so a group of sixteen costs eight times the sampling compute of a group of two for the same number of prompts. Move the slider and watch the readout's advantages settle: at two attempts they are pinned near plus and minus one, and by sixteen the outliers separate from the pack.
One prompt's sampled group on a reward axis, the group mean as the baseline, and each attempt's signed distance from that mean. The advantages in the readout come from Multimodal.rl.groupAdvantage.
Advantage = (reward minus group mean) / group spread. Small groups give a noisy baseline; large groups give a stable one and cost more samples per prompt.
Further reading
The staging recipe, the instruction-tuning data, and the methods that followed it. Read them in this order and the arc runs from the three-stage schedule, through the mixture and unfreezing decisions that refine it, to the objectives that teach a trained model to prefer a better answer, and finally to the reinforcement-learning stage that teaches it to check one.
- Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Jae Lee, "Visual Instruction Tuning", 2023 — LLaVA's two-stage recipe: projector warm-up, then instruction tuning.
- Haotian Liu, Chunyuan Li, Yuheng Li and Yong Jae Lee, "Improved Baselines with Visual Instruction Tuning", 2023 — the stage-two mixture, unfreezing decisions and the 665K instruction set.
- Qinghao Ye, Haiyang Xu, Guohai Xu and coauthors, "mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality", 2023 — staged training with a modular visual interface and an explicit forgetting discussion.
- Rafael Rafailov, Archit Sharma, Eric Mitchell and coauthors, "Direct Preference Optimization: Your Language Model is Secretly a Reward Model", 2023 — the objective that multimodal DPO adapts to image-conditioned pairs.
- DeepSeek-AI, Daya Guo, Dejian Yang and coauthors, "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning", 2025 — group-relative policy optimisation and the verifiable-reward recipe behind the reasoning wave.
Cheat sheet
| Term | Meaning here |
|---|---|
| Warm-up stage | Train the bridge alone against a frozen language model on caption data |
| Joint pretraining | Unfreeze the language model, train it with the bridge on a mixed multimodal corpus |
| Instruction tuning | Small, curated visual dialogues that turn a captioner into an assistant |
| Catastrophic forgetting | Text ability degrades as the language model is tuned on image-dominated data |
| Freezing schedule | Which components receive gradients in which stage; the schedule's real content |
| Data mixture | The per-stage blend of captions, VQA, OCR, grounding and text-only data |
| Grounding lag | Naming an object is easier than pointing at it, so grounding trains last and lags |
| Preference tuning | Train on chosen/rejected pairs rather than on a single reference answer |
| DPO | Direct preference objective; $\beta$ sets how far the policy may drift from the reference |
| RL post-training | Sample attempts, score them, and update toward the better ones; no reference answer |
| GRPO | Group-relative policy optimisation: advantage = (reward minus group mean) / group spread, no critic |
| Verifiable reward | A score a rule can compute — a checked answer, a matched box — instead of a human preference |
| Thinking VLM | A model that writes intermediate reasoning before answering, bought with RL and paid for in latency |