FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Penedo, Guilherme, Kydlíček, Hynek, Sabolčec, Vinko, Messmer, Bettina, Foroutan, Negar, Kargaran, Amir Hossein, Raffel, Colin, Jaggi, Martin, Von Werra, Leandro, Wolf, Thomas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915359609585664
author Penedo, Guilherme
Kydlíček, Hynek
Sabolčec, Vinko
Messmer, Bettina
Foroutan, Negar
Kargaran, Amir Hossein
Raffel, Colin
Jaggi, Martin
Von Werra, Leandro
Wolf, Thomas
author_facet Penedo, Guilherme
Kydlíček, Hynek
Sabolčec, Vinko
Messmer, Bettina
Foroutan, Negar
Kargaran, Amir Hossein
Raffel, Colin
Jaggi, Martin
Von Werra, Leandro
Wolf, Thomas
contents Pre-training state-of-the-art large language models (LLMs) requires vast amounts of clean and diverse text data. While the open development of large high-quality English pre-training datasets has seen substantial recent progress, training performant multilingual LLMs remains a challenge, in large part due to the inherent difficulty of tailoring filtering and deduplication pipelines to a large number of languages. In this work, we introduce a new pre-training dataset curation pipeline based on FineWeb that can be automatically adapted to support any language. We extensively ablate our pipeline design choices on a set of nine diverse languages, guided by a set of meaningful and informative evaluation tasks that were chosen through a novel selection process based on measurable criteria. Ultimately, we show that our pipeline can be used to create non-English corpora that produce more performant models than prior datasets. We additionally introduce a straightforward and principled approach to rebalance datasets that takes into consideration both duplication count and quality, providing an additional performance uplift. Finally, we scale our pipeline to over 1000 languages using almost 100 Common Crawl snapshots to produce FineWeb2, a new 20 terabyte (5 billion document) multilingual dataset which we release along with our pipeline, training, and evaluation codebases.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20920
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Penedo, Guilherme
Kydlíček, Hynek
Sabolčec, Vinko
Messmer, Bettina
Foroutan, Negar
Kargaran, Amir Hossein
Raffel, Colin
Jaggi, Martin
Von Werra, Leandro
Wolf, Thomas
Computation and Language
Pre-training state-of-the-art large language models (LLMs) requires vast amounts of clean and diverse text data. While the open development of large high-quality English pre-training datasets has seen substantial recent progress, training performant multilingual LLMs remains a challenge, in large part due to the inherent difficulty of tailoring filtering and deduplication pipelines to a large number of languages. In this work, we introduce a new pre-training dataset curation pipeline based on FineWeb that can be automatically adapted to support any language. We extensively ablate our pipeline design choices on a set of nine diverse languages, guided by a set of meaningful and informative evaluation tasks that were chosen through a novel selection process based on measurable criteria. Ultimately, we show that our pipeline can be used to create non-English corpora that produce more performant models than prior datasets. We additionally introduce a straightforward and principled approach to rebalance datasets that takes into consideration both duplication count and quality, providing an additional performance uplift. Finally, we scale our pipeline to over 1000 languages using almost 100 Common Crawl snapshots to produce FineWeb2, a new 20 terabyte (5 billion document) multilingual dataset which we release along with our pipeline, training, and evaluation codebases.
title FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
topic Computation and Language
url https://arxiv.org/abs/2506.20920