When not to fine-tune
Fine-tuning has become the most over-prescribed move in applied LLM work: it is reachable in an afternoon, it feels like real engineering, and it is usually the wrong first response to a failure. The disciplined order is prompt, then retrieval, then an eval you can trust, then — only if the failure is a behaviour the model cannot be instructed into — tuning, and finally reinforcement learning when the outcome is programmatically verifiable. This part is the decision tree that keeps you off the treadmill, and the reason the durable asset you are building is a labelled dataset, not a set of weights.
The decision tree
Prompt → RAG → eval → tune → RL, in that order
Every fine-tuning project should survive four questions before it starts. Can you measure the failure? If not, you have nothing to tune toward and nothing to prove you improved — build the eval first. Is it a knowledge gap? A model that does not know your private facts, your current prices, or your internal vocabulary is not misbehaving; it is uninformed, and the cheap fix is retrieval, not weights. Is it a behaviour, format or style problem? A model that knows the right answer but returns the wrong JSON shape, ignores a house style, or reasons in the wrong register is a genuine tuning candidate, because prompting is a soft lever and behaviour is what tuning hardens. Do you have labelled data? Without examples of the target output there is nothing to train on, and the work that produces them is the work that matters.
Toggle the gates below. The recommendation and its reasoning change with each state; the point is that fine-tuning only appears at the end of a chain, never at the start.
Four gates on the left, one recommendation on the right. An unmet gate routes you away from tuning.
Knowledge versus behaviour
Two axes: what it knows, and how it acts
The single most useful cut in this decision is between a knowledge deficiency and a behaviour deficiency. Knowledge is anything the model must be told: facts, documents, policy, current state. It changes constantly, it lives in your systems, and retrieval brings it in at request time with a citation. Behaviour is anything the model must be: a response schema, a tone, a refusal pattern, a tool-call policy, a decomposition habit. It is stable, it is expensive to express in a prompt, and it is exactly what examples teach.
The second axis is whether you can measure the failure at all. A knowledge gap you cannot evaluate is indistinguishable from a behaviour gap you cannot evaluate, and tuning against an unmeasured target is how teams ship regressions they cannot see. The quadrant below places six common failures; pick one to see the recommended fix and why.
Horizontal: knowledge deficit (left) to behaviour deficit (right). Vertical: measurable (top) to unmeasurable (bottom).
The upgrade treadmill
Every new base model invalidates the last fine-tune
A fine-tune is bound to the base checkpoint it was trained against. When the provider ships a better base model, your adapter does not come along: it must be retrained on the new base, and any quality you had is provisional until it is. That is the treadmill. Teams that treat the weights as the asset relabel, retrain and re-verify on every upgrade, and each cycle's work is discarded by the next. Teams that treat the labelled dataset and the eval as the asset keep the same examples and the same measurement, point them at the new base, and inherit the improvement instead of paying for it again.
The chart runs three base-model generations for the two strategies. The dataset strategy's quality compounds because its data and its eval improve with use; the weights-only strategy re-earns the same gain each time.
Seeded quality curves across three generations. Solid = refit each time; dashed = keep the dataset and re-evaluate.
What you actually own
The dataset appreciates; the weights depreciate
Run the accounting honestly and the two assets move in opposite directions. The fine-tuned weights are a snapshot of one base model at one moment: their value decays to nothing the day the base is retired, and it decays faster the faster the frontier moves. The labelled dataset and the eval harness do not decay. They become more valuable with every failure they encode, they transfer to whichever base model wins next, and they are the one artifact in the whole pipeline that a competitor cannot copy from your API. This is why "we fine-tuned a model" is a weaker claim than "we have ten thousand labelled cases and a harness that runs them in minutes."
Move the churn slider to see how fast the weights decay relative to the dataset's compounding value across three upgrades.
Relative value after each base-model generation. Weights start high and fall; dataset + eval start lower and climb.
When reinforcement learning genuinely applies
Verifiable rewards, reasoning policies, and a real bill
Above supervised tuning sits reinforcement learning from verifiable rewards (RLVR). It is not a general-purpose upgrade over fine-tuning; it applies when the outcome is programmatically checkable — a unit test passes, a tool call is valid, a proof checks, a game is won — and the task rewards a long policy of intermediate steps rather than a single output. That is why it dominates reasoning and tool-use work: a verifier can grade a thousand sampled trajectories without a human, and the policy learns a strategy rather than a style. Group-relative methods such as GRPO make this practical by comparing samples within a group instead of training a separate value model.
The bill is real: high sample counts, expensive rollouts, reward-model or verifier infrastructure, and the characteristic failure mode of reward hacking — the policy learns to satisfy the checker rather than the task. It is also the method most likely to be applied too early, to a problem whose real defect is a missing fact that RAG would have supplied for free.
Cheat sheet
| Question | The answer that shapes the build |
|---|---|
| Where does fine-tuning sit in the order? | Last, after prompt, retrieval and an eval you trust — and before RL, which is more expensive still. |
| The model lacks a fact I need. Tune it? | No. Retrieve the fact at request time; facts change and weights cannot. |
| The model ignores my output format. Tune it? | Yes — behaviour, format and style are exactly what examples harden. |
| I cannot measure the failure. What first? | Build the eval. Tuning without measurement ships regressions you cannot see. |
| What happens on the next base-model upgrade? | The fine-tune is invalidated and must be redone; the dataset and eval carry forward unchanged. |
| Which asset compounds over time? | The labelled dataset and the eval harness. The weights depreciate to zero when the base is retired. |
| When is RLVR the right tool? | When rewards are verifiable and the target is a reasoning or tool-use policy — never as a first response to a knowledge gap. |
| What is RLVR's signature failure? | Reward hacking: the policy exploits the verifier instead of solving the task. Hold out verifiers and read the traces. |
Further reading
- Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models", ICLR 2022 — why adapters are cheap and how a frozen base plus a low-rank update changes behaviour.
- Dettmers et al., "QLoRA: Efficient Finetuning of Quantized LLMs", NeurIPS 2023 — the single-GPU path this part's dataset argument still applies to.
- Shao et al., "DeepSeekMath", 2024 — introduced GRPO, the group-relative method behind much of the RLVR wave.
- DeepSeek-AI, "DeepSeek-R1", 2025 — verification-driven reinforcement learning of reasoning at scale, and its cost.
- Yao et al., "τ-bench", 2024 — tool-agent-user tasks with a verifiable end state, the kind of target RLVR can grade.
- Jimenez et al., "SWE-bench", ICLR 2024 — repository issue resolution graded by a test suite: a verifiable reward, and a contaminated public signal.