Building the pretraining dataset
A base model is, in a real sense, a compressed summary of the text it was trained on — so what goes into that text matters enormously, and gathering it is its own multi-stage engineering problem, well before any GPU runs a training step. We'll follow Dolma, Ai2's fully open pretraining corpus for OLMo, through the pipeline that turns raw web crawls into training-ready text — with real published proportions, not illustrative guesses.
Where the text comes from
Sourcing
No single source has enough high-quality text at trillion-token scale, so pretraining corpora blend several. Below are Dolma v1.6's actual published token proportions (Soldaini et al., 2024; Ai2's dataset card) — not an illustrative guess. Notice how dominant web crawl text is by volume, and how small (by percentage) the most carefully-curated sources are.
Click a bar for a one-line description. Source: Dolma v1.6, Soldaini et al. 2024 / allenai/dolma dataset card.
The filtering funnel
Cleaning
Raw Common Crawl is mostly not useful training text — boilerplate, navigation menus, spam, non-English pages, near-duplicate pages mirrored across sites. Dolma's pipeline (released as open-source code, not just a description) runs each document through a sequence of filters, each one discarding a further slice. Dolma's own paper reports the aggregate result: roughly 200 TB of raw crawled text is curated down to an ~11 TB corpus (about 3 trillion tokens) — a real, published endpoint. The intermediate per-stage percentages below are a representative shape (Dolma's paper doesn't break out a single universal per-stage retention table across all sources), but they're built to land on that real ~5–6% overall retention.
Quality-classifier threshold trade-off
A quality classifier outputs a score per document; you pick a cutoff. Raise it and you keep cleaner text but throw more of it away — precision vs. recall, made concrete with a small labelled set of 40 toy documents (half genuinely "high quality," half "low quality," scores drawn to overlap realistically).
Deduplication: MinHash and LSH banding
Cleaning
The web republishes the same content constantly — syndicated news, mirrored docs, boilerplate templates. Exact-duplicate removal (hashing whole documents) catches identical copies; Dolma also runs fuzzy deduplication using MinHash plus locality-sensitive hashing (LSH), which flags likely-similar document pairs without ever comparing every pair directly — infeasible at trillions of documents.
The similarity measure is Jaccard similarity between two documents' sets of overlapping word "shingles":
Comparing full shingle sets for every pair is still too slow at scale. MinHash + LSH compresses each document to a small signature, splits that signature into b bands of r rows each, and only checks pairs that collide in at least one band — turning an O(n²) all-pairs comparison into a fast lookup, at the cost of a probabilistic (not exact) similarity threshold, whose sharpness is tunable via b and r.
The steep S-curve is the point: with the right (b, r), pairs above the similarity threshold are caught with high probability while pairs below it are almost never even checked — that's what makes fuzzy dedup at trillion-token scale tractable at all.
Synthetic rewriting & quality scaling
Turning good text into more good text
Filtering throws text away; rewriting multiplies it. OLMo 3's midtraining mix is built substantially from synthetic re-creations of high-quality datasets that had restrictive licenses, plus new data generated around existing problems. The pattern is consistent: take text that is already good but awkward, sparse, or legally unusable, and use a capable open model to produce a larger, cleaner, permissively-licensed version. Every one of these is validated with a microanneal — a small 5–10B-token annealing run, compared against a web-only baseline — before it earns a place in the 100B mix.
- TinyMATH — for each of the 7,500 MATH training problems, generate 100 similar problems, plus a Python chain-of-thought solution (PoT) and conversational explanations (MIND): 1.14B tokens, +13.2 MATH / +13.9 GSM8K.
- CraneMath — an independent, permissive recreation of SwallowMath by rewriting FineMath4+ with Qwen3: 5.6B tokens, +18.5 MATH / +27.4 GSM8K.
- MegaMatt — MegaMath-Web-Pro-Max recreated from post-June-2023 MegaMath-Web with Qwen3 rewrites: 3.88B tokens, +8.0 MATH / +13.0 GSM8K.
- CraneCode — a permissive SwallowCode recreation from the-stack-v2-smol Python, filtered for syntax and lint errors: 18.8B tokens, +5.0 HumanEval.
- Educational-value scaling — Stack-Edu documents are bucketed by an educational-value classifier, then sampled with reservoir sampling weighted toward the upper 20% of buckets, which beat both the natural distribution and simpler top-per-language sampling. 50% of code documents are transformed into fill-in-the-middle (FIM) tasks.
Decontamination: keeping the exam out of the textbook
Cleaning
If a benchmark's evaluation questions (or worse, its answers) leak into pretraining data, the model can score well by memorization rather than genuine capability — inflating reported results. The same n-gram-overlap machinery from deduplication is reused here, just checked against benchmark text instead of other training documents.
The two-phase decon method
At benchmark scale, checking every n-gram of every training document against every eval question is too slow. OLMo 3's open-source decon tool splits the job into two phases. Detection samples n-grams at a regular stride — a cheap, non-overlapping scan that finds candidate matches fast. Cluster expansion then expands each match outward, counting how many adjacent n-grams are also contaminated, and only removes the document if that count clears a threshold. The two phases are what make it both fast and accurate: the stride skips most positions, and the expansion recovers the local detail needed to avoid false positives.
The subtlety is that contamination and overestimation are not the same thing. A benchmark can be heavily contaminated without its score being inflated. In OLMo 3's audit, DROP's generative metric showed a massive +13.99 points of overestimation when contaminated text was kept, but GSM8K was 84.99% contaminated and still did better after decontamination (78.49 → 80.10) — because the leaked format didn't match how the benchmark is actually scored. Minerva was 45% contaminated yet slightly worse with contamination (−1.95). High contamination matters most when the task is difficult, unsaturated, and formatted the same way it was leaked.
Toy contamination detector: documents scored for overlap, split into truly contaminated (top) and clean (bottom). The threshold trades precision against recall.
| Benchmark | % contaminated | Perf. overestimation (contaminated − decontaminated) |
|---|---|---|
| DROP (generative) | 66.87% | +13.99 |
| GSM-Symbolic | 0.00% | +5.54 |
| Codex HumanEval @16 | 7.07% | +3.54 |
| SQuAD | 83.24% | +1.70 |
| GSM8K | 84.99% | +1.61 (but decontaminated perf. higher) |
| Minerva | 45.17% | −1.95 |
Source: OLMo 3 report, arXiv:2512.13961v2, Tables 17–18.
Data mixing and epoch budgets
Wrap-up
The final step isn't just "keep what survives filtering" — it's choosing how much of each surviving source to actually show the model. Reweighting a small, high-quality source upward means training on it for more than one epoch (full pass) — fine in moderation, but repeating low-diversity text too many times can hurt more than it helps. Adjust the sliders below (starting from Dolma v1.6's real proportions) against a fixed 3-trillion-token training budget and watch the implied epoch count per source.
Cheat sheet
Recap
| Stage | Goal | Typical technique |
|---|---|---|
| Sourcing | Assemble diverse, large-scale raw text | Common Crawl + code, papers, books, encyclopedic text |
| Language filter | Keep only target-language text | Language ID classifiers |
| Quality filter | Discard boilerplate / low-value text | Heuristic rules + trained classifiers |
| Deduplication | Remove exact and near-duplicate documents | Exact hashing + MinHash/LSH banding |
| PII / safety scrub | Remove personal data and clearly unsafe content | Pattern matching + classifiers |
| Decontamination | Prevent eval-set leakage into training data | Overlap search against benchmark text |
| Synthetic rewriting | Scale and clean good text; work around licenses | LLM rewrites (CraneMath, CraneCode, MegaMatt, TinyMATH) |
| Educational-value scaling | Sample the most useful code/math | Quality/value bucketing + reservoir sampling; 50% FIM |
| Two-phase decon | Find leaked eval text fast and accurately | Stride detection + cluster expansion against benchmarks |
| Mixing | Weight sources for the best downstream model | Small-scale ablations before committing at scale |
Further reading
References
- Soldaini et al., "Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research" (2024).
- Broder, "On the resemblance and containment of documents" (1997) — MinHash.
- Penedo et al., "The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale" (2024).
- Dolma dataset card and code (Hugging Face / Ai2).