Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Chinchilla: the compute-optimal split

Scaling laws

Hoffmann et al. (DeepMind, 2022) trained hundreds of models at different sizes and token counts, then fit power-law curves to find how the loss-minimizing parameter count N and token count D each grow with compute C. All three of the paper's independent fitting methods agreed on the headline result to within a small factor: at compute-optimal allocation, N and D scale with roughly equal exponents in C, which empirically worked out to about 20 training tokens per parameter across the compute range they studied (their own three approaches predicted 67–68B/1.5T, 67B/1.5T, and 40B/1.9T at Gopher's compute budget — the exact fitted coefficients of the parametric loss curve are less robust than this ratio, which is why later reproduction attempts flagged ambiguity in the paper's reported constants but not in the ~20:1 conclusion).

$$C \approx 6ND, \qquad D_{\text{opt}} \approx 20\,N_{\text{opt}} \ \Rightarrow\ N_{\text{opt}} = \sqrt{\frac{C}{120}},\ \ D_{\text{opt}} = 20\,N_{\text{opt}}$$

Those two equations are exactly solvable — for a fixed compute budget C, there's a whole hyperbola of (N, D) pairs that cost exactly C FLOPs ($D = C/6N$); the compute-optimal point is where that hyperbola crosses the 20-tokens-per-parameter line. Drag the budget below and watch both curves and their intersection move.

⚠️ Why real models overtrain anyway: the compute-optimal point minimizes training compute for a target loss — it says nothing about inference cost. A smaller model trained on far more than its "optimal" token count (an overtrained model) reaches nearly the same loss but is cheaper to serve for the rest of its life, often billions of inference calls. OLMo 2 7B, like Llama 3 8B, is trained on several trillion tokens — tens of tokens per parameter past the Chinchilla-optimal ~20:1 ratio for its size.
2

From parameters and tokens to GPU-hours and dollars

Turning FLOPs into a training run

$C \approx 6ND$ is a rule of thumb — 2 FLOPs per parameter per token for the forward pass, ×3 to include the backward pass, giving 6 total. Divide by a GPU's real achieved throughput (never its advertised peak — see MFU below) to get GPU-hours, then by cluster size for wall-clock time.

3

Why a 70B model doesn't fit on one GPU

Memory accounting

Training needs more than just the weights: AdamW keeps two extra optimizer states (momentum and variance) per parameter, gradients need one more copy, and mixed-precision training typically keeps an fp32 master copy alongside a lower-precision (bf16/fp8) working copy. A rough per-parameter budget in mixed precision: 2 bytes (bf16 weights) + 2 bytes (bf16 grads) + 4+4 bytes (fp32 Adam moments) + 4 bytes (fp32 master weights) ≈ 16 bytes/parameter before activations — which is why a 70B model needs well over 1TB of GPU memory just for state, before it processes a single token.

4

Splitting a model across a cluster

Distributed training

When one model's state doesn't fit on one GPU (or you simply want more throughput), you shard the work. The four standard strategies compose:

Real training runs combine all four (often called "3D" or "4D parallelism") — e.g. FSDP within a node, pipeline parallel across node groups, data parallel across replicas of the whole pipeline. Communication cost, not memory, is usually the thing that limits how far you can push any single strategy — which is why interconnect bandwidth is as strategically important as GPU count.

✓

Cheat sheet

Recap

ConceptWhat it isRule of thumb
C ≈ 6NDTraining FLOPs ≈ 6 × parameters × tokens2 (fwd) + 4 (bwd) FLOPs per param per token
Compute-optimal ratioChinchilla's fitted loss-minimizing D/N~20 tokens per parameter
OvertrainingTraining well past compute-optimal to shrink inference costOLMo 2, Llama 3: 100s of tokens/param
Mixed-precision memoryWeights + grads + optimizer state + master copy~16 bytes/parameter (AdamW, bf16)
MFUModel FLOPs Utilization — achieved ÷ peak FLOP/sReal runs: 30–55%, rarely more
📚

Further reading

References

?

Check your understanding

0/5 answered
With a budget, a data split, and a cluster plan settled, it's time to actually run the training loop. Continue: pretraining, the training loop →