Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Where the text comes from

Sourcing

No single source has enough high-quality text at trillion-token scale, so pretraining corpora blend several. Below are Dolma v1.6's actual published token proportions (Soldaini et al., 2024; Ai2's dataset card) — not an illustrative guess. Notice how dominant web crawl text is by volume, and how small (by percentage) the most carefully-curated sources are.

Click a bar for a one-line description. Source: Dolma v1.6, Soldaini et al. 2024 / allenai/dolma dataset card.

Click a bar above.
💡 Why this matters: the mixture is itself a hyperparameter. Overweight web text and you get breadth but noise; overweight code and math ability improves but general fluency can suffer. Ai2 publishes not just the sources but the exact mixing ratios used for each Dolma release, so this trade-off is something you can actually inspect rather than guess at.
2

The filtering funnel

Cleaning

Raw Common Crawl is mostly not useful training text — boilerplate, navigation menus, spam, non-English pages, near-duplicate pages mirrored across sites. Dolma's pipeline (released as open-source code, not just a description) runs each document through a sequence of filters, each one discarding a further slice. Dolma's own paper reports the aggregate result: roughly 200 TB of raw crawled text is curated down to an ~11 TB corpus (about 3 trillion tokens) — a real, published endpoint. The intermediate per-stage percentages below are a representative shape (Dolma's paper doesn't break out a single universal per-stage retention table across all sources), but they're built to land on that real ~5–6% overall retention.

raw crawl → language filter → quality filter → deduplication → PII / toxic-content removal → training-ready text
⚠️ Quality filtering is the trickiest stage: Dolma uses a mix of heuristic rules (too short, too repetitive, too many symbols relative to words) and trained classifiers (does this page resemble the kind of text people actually found useful, e.g. Wikipedia-referenced pages) to score documents. Get the threshold wrong and you either keep boilerplate or throw away legitimate but unusual text — there's no ground truth to check against.

Quality-classifier threshold trade-off

A quality classifier outputs a score per document; you pick a cutoff. Raise it and you keep cleaner text but throw more of it away — precision vs. recall, made concrete with a small labelled set of 40 toy documents (half genuinely "high quality," half "low quality," scores drawn to overlap realistically).

3

Deduplication: MinHash and LSH banding

Cleaning

The web republishes the same content constantly — syndicated news, mirrored docs, boilerplate templates. Exact-duplicate removal (hashing whole documents) catches identical copies; Dolma also runs fuzzy deduplication using MinHash plus locality-sensitive hashing (LSH), which flags likely-similar document pairs without ever comparing every pair directly — infeasible at trillions of documents.

The similarity measure is Jaccard similarity between two documents' sets of overlapping word "shingles":

$$J(A,B) = \frac{|A \cap B|}{|A \cup B|}$$

Comparing full shingle sets for every pair is still too slow at scale. MinHash + LSH compresses each document to a small signature, splits that signature into b bands of r rows each, and only checks pairs that collide in at least one band — turning an O(n²) all-pairs comparison into a fast lookup, at the cost of a probabilistic (not exact) similarity threshold, whose sharpness is tunable via b and r.

$$P(\text{flagged as candidate pair} \mid \text{true similarity } s) = 1-(1-s^r)^b$$

The steep S-curve is the point: with the right (b, r), pairs above the similarity threshold are caught with high probability while pairs below it are almost never even checked — that's what makes fuzzy dedup at trillion-token scale tractable at all.

4

Synthetic rewriting & quality scaling

Turning good text into more good text

Filtering throws text away; rewriting multiplies it. OLMo 3's midtraining mix is built substantially from synthetic re-creations of high-quality datasets that had restrictive licenses, plus new data generated around existing problems. The pattern is consistent: take text that is already good but awkward, sparse, or legally unusable, and use a capable open model to produce a larger, cleaner, permissively-licensed version. Every one of these is validated with a microanneal — a small 5–10B-token annealing run, compared against a web-only baseline — before it earns a place in the 100B mix.

  • TinyMATH — for each of the 7,500 MATH training problems, generate 100 similar problems, plus a Python chain-of-thought solution (PoT) and conversational explanations (MIND): 1.14B tokens, +13.2 MATH / +13.9 GSM8K.
  • CraneMath — an independent, permissive recreation of SwallowMath by rewriting FineMath4+ with Qwen3: 5.6B tokens, +18.5 MATH / +27.4 GSM8K.
  • MegaMatt — MegaMath-Web-Pro-Max recreated from post-June-2023 MegaMath-Web with Qwen3 rewrites: 3.88B tokens, +8.0 MATH / +13.0 GSM8K.
  • CraneCode — a permissive SwallowCode recreation from the-stack-v2-smol Python, filtered for syntax and lint errors: 18.8B tokens, +5.0 HumanEval.
  • Educational-value scaling — Stack-Edu documents are bucketed by an educational-value classifier, then sampled with reservoir sampling weighted toward the upper 20% of buckets, which beat both the natural distribution and simpler top-per-language sampling. 50% of code documents are transformed into fill-in-the-middle (FIM) tasks.
💡 Why rewrite instead of just using the original? Licensing is one reason, scale is another: a 3.6B-token dataset can become 5.6B cleaner tokens, and the rewrite can normalize notation and add explicit reasoning that the original web text never had. The microanneal is what keeps this honest — synthetic data that doesn't move a benchmark doesn't ship.
5

Decontamination: keeping the exam out of the textbook

Cleaning

If a benchmark's evaluation questions (or worse, its answers) leak into pretraining data, the model can score well by memorization rather than genuine capability — inflating reported results. The same n-gram-overlap machinery from deduplication is reused here, just checked against benchmark text instead of other training documents.

💡 PII and safety scrubbing happen around the same stage: filters look for patterns resembling personal data (emails, phone numbers, ID numbers) and remove or mask them, and classifiers flag clearly toxic or unsafe pages for removal — separate from (and a preview of) the more targeted safety work in Part 12.

The two-phase decon method

At benchmark scale, checking every n-gram of every training document against every eval question is too slow. OLMo 3's open-source decon tool splits the job into two phases. Detection samples n-grams at a regular stride — a cheap, non-overlapping scan that finds candidate matches fast. Cluster expansion then expands each match outward, counting how many adjacent n-grams are also contaminated, and only removes the document if that count clears a threshold. The two phases are what make it both fast and accurate: the stride skips most positions, and the expansion recovers the local detail needed to avoid false positives.

The subtlety is that contamination and overestimation are not the same thing. A benchmark can be heavily contaminated without its score being inflated. In OLMo 3's audit, DROP's generative metric showed a massive +13.99 points of overestimation when contaminated text was kept, but GSM8K was 84.99% contaminated and still did better after decontamination (78.49 → 80.10) — because the leaked format didn't match how the benchmark is actually scored. Minerva was 45% contaminated yet slightly worse with contamination (−1.95). High contamination matters most when the task is difficult, unsaturated, and formatted the same way it was leaked.

Toy contamination detector: documents scored for overlap, split into truly contaminated (top) and clean (bottom). The threshold trades precision against recall.

Benchmark% contaminatedPerf. overestimation (contaminated − decontaminated)
DROP (generative)66.87%+13.99
GSM-Symbolic0.00%+5.54
Codex HumanEval @167.07%+3.54
SQuAD83.24%+1.70
GSM8K84.99%+1.61 (but decontaminated perf. higher)
Minerva45.17%−1.95

Source: OLMo 3 report, arXiv:2512.13961v2, Tables 17–18.

6

Data mixing and epoch budgets

Wrap-up

The final step isn't just "keep what survives filtering" — it's choosing how much of each surviving source to actually show the model. Reweighting a small, high-quality source upward means training on it for more than one epoch (full pass) — fine in moderation, but repeating low-diversity text too many times can hurt more than it helps. Adjust the sliders below (starting from Dolma v1.6's real proportions) against a fixed 3-trillion-token training budget and watch the implied epoch count per source.

✓

Cheat sheet

Recap

StageGoalTypical technique
SourcingAssemble diverse, large-scale raw textCommon Crawl + code, papers, books, encyclopedic text
Language filterKeep only target-language textLanguage ID classifiers
Quality filterDiscard boilerplate / low-value textHeuristic rules + trained classifiers
DeduplicationRemove exact and near-duplicate documentsExact hashing + MinHash/LSH banding
PII / safety scrubRemove personal data and clearly unsafe contentPattern matching + classifiers
DecontaminationPrevent eval-set leakage into training dataOverlap search against benchmark text
Synthetic rewritingScale and clean good text; work around licensesLLM rewrites (CraneMath, CraneCode, MegaMatt, TinyMATH)
Educational-value scalingSample the most useful code/mathQuality/value bucketing + reservoir sampling; 50% FIM
Two-phase deconFind leaked eval text fast and accuratelyStride detection + cluster expansion against benchmarks
MixingWeight sources for the best downstream modelSmall-scale ablations before committing at scale
📚

Further reading

References

?

Check your understanding

0/5 answered
Once the dataset exists, the actual training loop — turning trillions of tokens into a base model — has its own moving parts: loss, optimizers, schedules, and scale. Continue: pretraining, the training loop →