Scaling Laws & Training Systems
Two questions every training run has to answer before it starts: given a fixed compute budget, what's the best split between model size and data (scaling laws)? And given that split, how do you actually run it across a cluster without running out of memory or bandwidth (training systems)? Both have real, quantitative answers.
Chinchilla: the compute-optimal split
Scaling laws
Hoffmann et al. (DeepMind, 2022) trained hundreds of models at different sizes and token counts, then fit power-law curves to find how the loss-minimizing parameter count N and token count D each grow with compute C. All three of the paper's independent fitting methods agreed on the headline result to within a small factor: at compute-optimal allocation, N and D scale with roughly equal exponents in C, which empirically worked out to about 20 training tokens per parameter across the compute range they studied (their own three approaches predicted 67–68B/1.5T, 67B/1.5T, and 40B/1.9T at Gopher's compute budget — the exact fitted coefficients of the parametric loss curve are less robust than this ratio, which is why later reproduction attempts flagged ambiguity in the paper's reported constants but not in the ~20:1 conclusion).
Those two equations are exactly solvable — for a fixed compute budget C, there's a whole hyperbola of (N, D) pairs that cost exactly C FLOPs ($D = C/6N$); the compute-optimal point is where that hyperbola crosses the 20-tokens-per-parameter line. Drag the budget below and watch both curves and their intersection move.
From parameters and tokens to GPU-hours and dollars
Turning FLOPs into a training run
$C \approx 6ND$ is a rule of thumb — 2 FLOPs per parameter per token for the forward pass, ×3 to include the backward pass, giving 6 total. Divide by a GPU's real achieved throughput (never its advertised peak — see MFU below) to get GPU-hours, then by cluster size for wall-clock time.
Why a 70B model doesn't fit on one GPU
Memory accounting
Training needs more than just the weights: AdamW keeps two extra optimizer states (momentum and variance) per parameter, gradients need one more copy, and mixed-precision training typically keeps an fp32 master copy alongside a lower-precision (bf16/fp8) working copy. A rough per-parameter budget in mixed precision: 2 bytes (bf16 weights) + 2 bytes (bf16 grads) + 4+4 bytes (fp32 Adam moments) + 4 bytes (fp32 master weights) ≈ 16 bytes/parameter before activations — which is why a 70B model needs well over 1TB of GPU memory just for state, before it processes a single token.
Splitting a model across a cluster
Distributed training
When one model's state doesn't fit on one GPU (or you simply want more throughput), you shard the work. The four standard strategies compose:
Real training runs combine all four (often called "3D" or "4D parallelism") — e.g. FSDP within a node, pipeline parallel across node groups, data parallel across replicas of the whole pipeline. Communication cost, not memory, is usually the thing that limits how far you can push any single strategy — which is why interconnect bandwidth is as strategically important as GPU count.
Cheat sheet
Recap
| Concept | What it is | Rule of thumb |
|---|---|---|
| C ≈ 6ND | Training FLOPs ≈ 6 × parameters × tokens | 2 (fwd) + 4 (bwd) FLOPs per param per token |
| Compute-optimal ratio | Chinchilla's fitted loss-minimizing D/N | ~20 tokens per parameter |
| Overtraining | Training well past compute-optimal to shrink inference cost | OLMo 2, Llama 3: 100s of tokens/param |
| Mixed-precision memory | Weights + grads + optimizer state + master copy | ~16 bytes/parameter (AdamW, bf16) |
| MFU | Model FLOPs Utilization — achieved ÷ peak FLOP/s | Real runs: 30–55%, rarely more |
Further reading
References
- Hoffmann et al. (DeepMind), "Training Compute-Optimal Large Language Models" (2022) — Chinchilla.
- Kaplan et al. (OpenAI), "Scaling Laws for Neural Language Models" (2020) — the earlier scaling-law study Chinchilla revised.
- Rajbhandari et al., "ZeRO: Memory Optimizations Toward Training Trillion Parameter Models" (2019).
- OLMo 2 Team (Ai2), "2 OLMo 2 Furious" — published hardware logs and token counts.