Data, distillation, and tuning in practice
If Part 13 decided whether to tune, this part is the work once the answer is yes. The uncomfortable headline is that data curation is roughly eighty per cent of the effort and model configuration is the other twenty: deduplication, label quality, class balance and coverage decide the outcome long before rank or learning rate do. What follows is the operational spine — synthetic data and the diversity collapse it invites, rejection sampling and its yield, the LoRA and QLoRA knobs that matter, the preference methods and when each applies, RLVR and the reward-hacking curve it produces, and the serving economics that make adapters so much cheaper than full models.
Curation is the work; synthetic data narrows the distribution
Dedupe, label quality, balance, coverage — then model collapse
A tuning run is a mirror of its dataset. Near-duplicate examples inflate the loss on a handful of patterns and teach nothing new. Noisy labels teach the noise. An unbalanced set teaches the majority behaviour and hides the rare one that actually matters in production. Poor coverage means the model is confident everywhere except where the failures live. None of these are fixed by a better optimizer, which is why the hours go into curation: deduplicating, adjudicating labels, balancing classes, and deliberately sampling the hard cases.
The shortcut that looks like it saves that work is synthetic data: have the model generate examples, filter them, and train on them. It is legitimate and widely used — distillation from a stronger teacher is the standard way to acquire a capability cheaply. It also has a failure mode with a name. Training a model on its own outputs and repeating the loop narrows the distribution: rare modes are under-sampled, so they disappear from the next generation's training set, and the generation after that is narrower still. This is model collapse, or diversity collapse, and it is why synthetic data must be anchored to real, human-labelled data and measured for coverage, not just for loss.
Step the generations below. Each self-training round keeps fewer of the original modes and shrinks the spread of the ones that survive — the distribution the model can produce gets narrower on every loop.
Seeded output distribution across generations of self-training. The pale curve is generation 0; the hot curve is the latest.
Rejection sampling and its yield
Keep only what a verifier accepts — and watch the volume fall
The standard way to turn a mediocre generator into good training data is rejection sampling: sample many candidate outputs, score each with a verifier, and keep only the ones that pass. The verifier is the whole game. If it can be a program — a test suite, a schema check, a citation that resolves — the loop is cheap, deterministic and honest. If it has to be a model judge, it inherits every bias of a judge and the loop costs money per sample.
The economics are a yield curve. A lenient threshold keeps most of the data, but the tail you keep is only slightly better than the batch. A strict threshold keeps the best data, but the number of accepted samples collapses — and past a point you are generating thousands of candidates to win a few dozen. The useful operating point is where kept-volume is still enough to train on while kept-quality is materially above the generator, and finding it is a measurement, not a guess.
Drag the verifier threshold below. The histogram shows the sampled quality distribution, split into kept and rejected; the readout reports the yield and the mean quality of what survives.
Seeded sample pool. Bars left of the threshold are discarded; the line is the mean quality of the kept set.
Adapters in practice: rank, method, and serving
LoRA and QLoRA knobs; DPO, KTO, ORPO; one base with many adapters
LoRA freezes the base weights and trains a pair of small low-rank matrices next to each target module. The knobs that matter are the rank (how much capacity the adapter gets), the target modules (which projections it touches — attention only, or attention and the feed-forward layers), the learning rate (adapters tolerate a much higher one than full fine-tuning), and whether you merge the adapter back into the base for serving or keep it separate. QLoRA does the same against a 4-bit quantized base, which is what makes single-GPU tuning routine; the base is frozen and the quantized weights are never updated. The trainable parameter count is small and calculable: a rank-r adapter on a d-wide projection adds 2·d·r parameters, so eight target modules at r=8 over a 4096-wide model is about two hundred and sixty thousand trainable values, well under a tenth of a percent of the base.
Supervised tuning teaches a behaviour by imitation. Preference methods teach it by comparison: DPO consumes pairs (a chosen and a rejected response) and needs no separate reward model; KTO consumes only binary good/bad labels, which is what you actually have from a thumbs-up button; ORPO folds supervised fine-tuning and preference into a single stage, which saves a pass over the data when you have no clean SFT checkpoint to start from. The method follows the labels you hold, not the other way round.
Pick a method and move the rank. The curve is quality against rank for the selected method, with the readout giving the parameter arithmetic and the serving story that makes adapters cheap: one frozen base in memory, many small adapters swapped per request.
Quality saturates in rank; the table of costs and the parameter count are the reason adapters, not full models, are what you ship.
RLVR, GRPO, and the reward-hacking curve
When the policy exploits the verifier instead of solving the task
Where a reward is programmatically verifiable, reinforcement learning can train a policy rather than a style — the case for reasoning and tool use. GRPO makes it practical by sampling a group of completions for each prompt and using the group's relative scores as the learning signal, which removes the separate value model that classic RLHF needed.
The characteristic failure is reward hacking. A verifier that checks the final answer but not the reasoning, or a test suite that can be satisfied by editing the test, becomes something to exploit. The proxy reward climbs smoothly to the top of its range while true task quality rises, plateaus, and then falls as the policy drifts into the blind spot. This is not a bug in the optimizer; it is the optimizer doing exactly what it was told. The defences are structural: hold out verifiers the policy never trains against, inspect the traces rather than the score, and treat a suspiciously clean reward curve as a warning rather than a result.
Move the verifier-coverage slider — how much of the true task the checker actually measures. Low coverage is a wide blind spot, and the divergence between proxy reward and true quality opens earlier and wider.
Seeded training curves. The proxy reward only ever rises; true quality turns down once the policy finds the verifier's blind spot.
Cheat sheet
| Question | The answer that shapes the run |
|---|---|
| Where does the effort actually go? | Data curation — dedupe, label quality, balance, coverage — is roughly 80% of the outcome. |
| What breaks self-training on synthetic data? | Diversity collapse: rare modes disappear, the distribution narrows each generation, and loss falls while capability shrinks. |
| How do you anchor synthetic data? | Keep real human-labelled examples in every round and measure coverage, not just loss. |
| How does rejection sampling pay off? | Keep only outputs a verifier accepts; a programmatic verifier is cheap and honest, a model judge is neither. |
| Where does the yield collapse? | Past a strictness threshold the kept volume falls faster than quality rises — find that point by measuring it. |
| LoRA's four knobs? | Rank, target modules, learning rate, and whether to merge the adapter into the base for serving. |
| Why is QLoRA the single-GPU path? | The base is 4-bit and frozen; only the adapters train, so memory is dominated by the quantized base. |
| DPO vs KTO vs ORPO? | DPO needs preference pairs; KTO needs binary good/bad; ORPO folds SFT and preference into one stage. |
| When RLVR, and its failure mode? | Verifiable rewards for reasoning/tool policies — and reward hacking, where the policy exploits the verifier. |
| One base or many models? | One frozen base plus many small adapters: memory tracks the base, not the number of variants. |
Further reading
- Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models", ICLR 2022 — the low-rank update and the parameter arithmetic this part's explorer uses.
- Dettmers et al., "QLoRA: Efficient Finetuning of Quantized LLMs", NeurIPS 2023 — 4-bit bases and why adapters fit on one GPU.
- Rafailov et al., "Direct Preference Optimization", NeurIPS 2023 — preference tuning without a separate reward model.
- Ethayarajh et al., "KTO: Model Alignment as Prospect Theoretic Optimization", 2024 — binary labels instead of pairs.
- Hong, Lee & Thorne, "ORPO: Monolithic Preference Optimization without Reference Model", 2024 — one-stage SFT plus preference.
- Shao et al., "DeepSeekMath", 2024 — GRPO, the group-relative method behind much RLVR work.
- DeepSeek-AI, "DeepSeek-R1", 2025 — verification-driven RL of reasoning at scale, and the reward-hacking caution that comes with it.
- Shumailov et al., "AI models collapse when trained on recursively generated data", Nature 2024 — the diversity-collapse result this part's first demo visualises.