Supervised Fine-Tuning & PEFT
A base model straight out of pretraining is a very good text completer, not an assistant. Ask it a question and it may just continue the pattern — writing more questions, or drifting into unrelated text — because that's the most statistically plausible continuation of internet-style text. Supervised fine-tuning (SFT) is the (comparatively small, but behaviorally huge) second training stage that reshapes this into "answer directly, follow instructions, stop when done." This part also covers how to do that fine-tuning cheaply, with LoRA.
Base model vs. instruction-tuned model
The behavior shift
Same prompt, same underlying knowledge from pretraining — very different behavior after SFT. These example completions are illustrative (written for this page, not sampled from a live model), but they show the qualitative shift accurately.
Instruction-tuning data: curation, synthesis, and distillation
What SFT trains on
SFT data is a curated set of (prompt, response) pairs — ideally covering a wide range of task types and response styles. For OLMo, Ai2 uses and publishes the Tulu mixture: a combination of human-written examples, carefully filtered existing instruction datasets, and synthetically generated examples — a real mix of curation and distillation (generating training examples from a stronger existing model's outputs, then filtering the results), rather than either alone.
SFT hyperparameters reflect this: learning rates are typically ~100× smaller than peak pretraining rates (the model is already good at language; SFT should nudge, not overwrite), and training usually runs only 2–3 epochs over the data, since with so few examples relative to pretraining, more epochs risks the overfitting-to-surface-patterns problem below.
Chat templates and multi-turn masking
Mechanics
A multi-turn conversation has to be serialized into a single flat token sequence, using special tokens to mark role boundaries (exact tokens vary by model family). SFT reuses the exact cross-entropy loss from Part 5 — the only change is which tokens count toward it: loss is masked on the system prompt and every user turn, and computed only on assistant-response tokens, across however many turns the conversation has.
Build a small multi-turn conversation below and see it formatted, with the mask applied per turn:
highlighted = contributes to the SFT loss (response tokens, both turns). Faded = present in the sequence, masked out of the loss.
LoRA: fine-tuning without touching most of the weights
Parameter-efficient fine-tuning
Full fine-tuning updates every one of a model's weight matrices, which for a 7B+ model means storing optimizer state for billions of parameters — expensive, and every fine-tune produces a full-size copy of the model. LoRA (Low-Rank Adaptation) instead freezes the original weight matrix W entirely and learns a small additive update as a product of two thin matrices, $\Delta W = BA$, where B is d×r and A is r×d for a small rank r (often 8–64) — far fewer trainable parameters than d×d.
Below, a toy 8×8 weight matrix has a "true" full-fine-tune update it would need to learn (a fixed, moderately complex target). Drag the rank slider and watch LoRA's B·A approximate it via real gradient descent — trainable-parameter count and reconstruction error both update live.
QLoRA layers one more idea on top: keep the frozen base weights W quantized to 4 bits (Part 11 covers quantization) while the small LoRA matrices train in higher precision — cutting base-model memory further, at the cost of a little extra quantization error on top of LoRA's own approximation error.
| Full fine-tuning | LoRA | QLoRA | |
|---|---|---|---|
| Trainable params (7B model, typical) | ~7B (100%) | ~0.1–1% | ~0.1–1% |
| Base weights | Updated directly | Frozen, bf16 | Frozen, 4-bit |
| Optimizer memory | Full AdamW state over all params | Only over LoRA matrices | Only over LoRA matrices |
| Output artifact | A new full-size model copy | A small adapter file (MBs) on top of the base | A small adapter file on top of the quantized base |
What SFT changes, and the risk of forgetting
Scope
SFT's effect is mostly on behavior and format, not the underlying knowledge set during pretraining on vastly more data. Two related risks: over-training on a narrow SFT set can cause the model to overfit to surface patterns (a fixed response length, a signature phrase); and fine-tuning too aggressively on a narrow domain can cause catastrophic forgetting, where capabilities the base model had — general knowledge, other task types — measurably degrade, sometimes called the alignment tax when the specific capability lost is general-purpose quality traded for narrower instruction-following. LoRA's frozen base weights are one structural mitigation: since most of the network is literally untouched, broad pretrained knowledge is mechanically harder to overwrite than with full fine-tuning.
Specialized SFT: tools and the reasoning-first warm start
OLMo 3 Instruct
SFT is not one homogeneous stage. OLMo 3 trains two different assistants from the same base, and each adds its own SFT flavor. Function-calling SFT teaches the model to emit structured tool calls: the training mix combines real trajectories executed against MCP servers (the Asta Scientific Corpus and Serper) with the much larger LLM-simulated SimFC dataset. The single most important lesson from that work is that the format has to be unified across every dataset — OpenAPI tool specs, pythonic code-block calls wrapped in XML tags, a dedicated environment role, and (for Instruct) tool-specific special tokens. Mixing conventions is a direct cause of unreliable tool use. Part 14 has the full format and the evaluation harness.
The reasoning-first warm start is the other notable choice. Rather than teaching instruction-following from the base model directly, OLMo 3 initializes the Instruct SFT stage from the already-trained Think SFT checkpoint. The instruction model then learns to suppress the thinking traces and answer concisely, but it inherits the reasoning skill underneath. The gain is broad and real, without making responses longer:
| Instruct SFT starting point | Avg. | GPQA | MATH | GSM8K | OMEGA | MBPP | IFEval |
|---|---|---|---|---|---|---|---|
| No thinking SFT first | 44.5 | 29.7 | 60.3 | 87.6 | 8.6 | 54.1 | 81.0 |
| With thinking SFT first (OLMo 3) | 47.8 | 34.4 | 65.9 | 91.1 | 12.2 | 57.1 | 84.7 |
| Gain | +3.3 | +4.7 | +5.6 | +3.5 | +3.6 | +3.0 | +3.7 |
Source: OLMo 3 report, arXiv:2512.13961v2, Table 13 (intermediate OLMo 3 Instruct 7B).
Cheat sheet
Recap
| Aspect | Pretraining | SFT |
|---|---|---|
| Objective | Predict next token, everywhere | Predict next token, only on response spans |
| Data scale | Trillions of tokens | Thousands–millions of examples |
| Learning rate | Peak ~1e-3 to 1e-4 scale | ~100× smaller |
| What it teaches | Language, facts, world knowledge | Format, directness, instruction-following behavior |
| LoRA | — | Freezes base weights, trains a low-rank BA update instead |
Further reading
References
- Ouyang et al., "Training Language Models to Follow Instructions with Human Feedback" (InstructGPT, 2022).
- Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models" (2021).
- Dettmers et al., "QLoRA: Efficient Finetuning of Quantized LLMs" (2023).
- Lambert et al. (Ai2), "Tulu 3" (2024) — SFT data mixture and recipe.