Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

The decision tree

Prompt → RAG → eval → tune → RL, in that order

Every fine-tuning project should survive four questions before it starts. Can you measure the failure? If not, you have nothing to tune toward and nothing to prove you improved — build the eval first. Is it a knowledge gap? A model that does not know your private facts, your current prices, or your internal vocabulary is not misbehaving; it is uninformed, and the cheap fix is retrieval, not weights. Is it a behaviour, format or style problem? A model that knows the right answer but returns the wrong JSON shape, ignores a house style, or reasons in the wrong register is a genuine tuning candidate, because prompting is a soft lever and behaviour is what tuning hardens. Do you have labelled data? Without examples of the target output there is nothing to train on, and the work that produces them is the work that matters.

Toggle the gates below. The recommendation and its reasoning change with each state; the point is that fine-tuning only appears at the end of a chain, never at the start.

Four gates on the left, one recommendation on the right. An unmet gate routes you away from tuning.

💡 The durable idea: fine-tuning changes how the model responds, not what it knows. If your failure is a missing fact, no amount of training fixes it; if it is a missing format, only training makes it stick.
2

Knowledge versus behaviour

Two axes: what it knows, and how it acts

The single most useful cut in this decision is between a knowledge deficiency and a behaviour deficiency. Knowledge is anything the model must be told: facts, documents, policy, current state. It changes constantly, it lives in your systems, and retrieval brings it in at request time with a citation. Behaviour is anything the model must be: a response schema, a tone, a refusal pattern, a tool-call policy, a decomposition habit. It is stable, it is expensive to express in a prompt, and it is exactly what examples teach.

The second axis is whether you can measure the failure at all. A knowledge gap you cannot evaluate is indistinguishable from a behaviour gap you cannot evaluate, and tuning against an unmeasured target is how teams ship regressions they cannot see. The quadrant below places six common failures; pick one to see the recommended fix and why.

Horizontal: knowledge deficit (left) to behaviour deficit (right). Vertical: measurable (top) to unmeasurable (bottom).

⚠️ The trap: "the model got it wrong" is not a diagnosis. Until you can say whether the missing ingredient was a fact or an action, you are choosing a fix by mood.
3

The upgrade treadmill

Every new base model invalidates the last fine-tune

A fine-tune is bound to the base checkpoint it was trained against. When the provider ships a better base model, your adapter does not come along: it must be retrained on the new base, and any quality you had is provisional until it is. That is the treadmill. Teams that treat the weights as the asset relabel, retrain and re-verify on every upgrade, and each cycle's work is discarded by the next. Teams that treat the labelled dataset and the eval as the asset keep the same examples and the same measurement, point them at the new base, and inherit the improvement instead of paying for it again.

The chart runs three base-model generations for the two strategies. The dataset strategy's quality compounds because its data and its eval improve with use; the weights-only strategy re-earns the same gain each time.

Seeded quality curves across three generations. Solid = refit each time; dashed = keep the dataset and re-evaluate.

4

What you actually own

The dataset appreciates; the weights depreciate

Run the accounting honestly and the two assets move in opposite directions. The fine-tuned weights are a snapshot of one base model at one moment: their value decays to nothing the day the base is retired, and it decays faster the faster the frontier moves. The labelled dataset and the eval harness do not decay. They become more valuable with every failure they encode, they transfer to whichever base model wins next, and they are the one artifact in the whole pipeline that a competitor cannot copy from your API. This is why "we fine-tuned a model" is a weaker claim than "we have ten thousand labelled cases and a harness that runs them in minutes."

Move the churn slider to see how fast the weights decay relative to the dataset's compounding value across three upgrades.

Relative value after each base-model generation. Weights start high and fall; dataset + eval start lower and climb.

5

When reinforcement learning genuinely applies

Verifiable rewards, reasoning policies, and a real bill

Above supervised tuning sits reinforcement learning from verifiable rewards (RLVR). It is not a general-purpose upgrade over fine-tuning; it applies when the outcome is programmatically checkable — a unit test passes, a tool call is valid, a proof checks, a game is won — and the task rewards a long policy of intermediate steps rather than a single output. That is why it dominates reasoning and tool-use work: a verifier can grade a thousand sampled trajectories without a human, and the policy learns a strategy rather than a style. Group-relative methods such as GRPO make this practical by comparing samples within a group instead of training a separate value model.

The bill is real: high sample counts, expensive rollouts, reward-model or verifier infrastructure, and the characteristic failure mode of reward hacking — the policy learns to satisfy the checker rather than the task. It is also the method most likely to be applied too early, to a problem whose real defect is a missing fact that RAG would have supplied for free.

💡 Order of operations: prompt → retrieval → eval → supervised tune → RL. Skip a rung only when you can name, in writing, why it cannot help — and remember that the rung you skip is usually the one that would have told you whether the next one worked.

Cheat sheet

QuestionThe answer that shapes the build
Where does fine-tuning sit in the order?Last, after prompt, retrieval and an eval you trust — and before RL, which is more expensive still.
The model lacks a fact I need. Tune it?No. Retrieve the fact at request time; facts change and weights cannot.
The model ignores my output format. Tune it?Yes — behaviour, format and style are exactly what examples harden.
I cannot measure the failure. What first?Build the eval. Tuning without measurement ships regressions you cannot see.
What happens on the next base-model upgrade?The fine-tune is invalidated and must be redone; the dataset and eval carry forward unchanged.
Which asset compounds over time?The labelled dataset and the eval harness. The weights depreciate to zero when the base is retired.
When is RLVR the right tool?When rewards are verifiable and the target is a reasoning or tool-use policy — never as a first response to a knowledge gap.
What is RLVR's signature failure?Reward hacking: the policy exploits the verifier instead of solving the task. Hold out verifiers and read the traces.

Further reading

6

Check your understanding

0/5 answered