FineWeb
Free while signed in. Answers cite the passages they came from.

HuggingFace's FineWeb is a 15 trillion token English web dataset built from 96 CommonCrawl snapshots (2013-2024). In 1.8B-parameter ablations, models trained on FineWeb beat C4, RefinedWeb, Dolma, The Pile, SlimPajama, and RedPajama2 across aggregated benchmarks.
Scale: 15T tokens spanning 52.5B documents (~50TB on disk), with released sample subsets at 10B, 100B, and 350B tokens for smaller experiments.
Filtering pipeline: Built on the open-source datatrove library - URL blocklists, Trafilatura text extraction, FastText English filter (>0.65), Gopher + C4 quality filters, custom FineWeb heuristics, MinHash deduplication, and PII anonymization.
Per-dump deduplication insight: Ablations show per-dump MinHash dedup beats global dedup, a non-obvious finding that shaped the final pipeline and contradicts a common assumption that more aggressive dedup is always better.
Fully reproducible: Released under ODC-By 1.0 with the complete datatrove pipeline and all ablation model checkpoints public, making it one of the most transparent large-scale web corpora to date.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack