FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Fuente:
arXiv
Saved in:
| Main Authors: | Penedo, Guilherme, Kydlíček, Hynek, Sabolčec, Vinko, Messmer, Bettina, Foroutan, Negar, Kargaran, Amir Hossein, Raffel, Colin, Jaggi, Martin, Von Werra, Leandro, Wolf, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
by: Penedo, Guilherme, et al.
Published: (2024)
by: Penedo, Guilherme, et al.
Published: (2024)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025)
by: Messmer, Bettina, et al.
Published: (2025)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026)
by: Turki, Yassine, et al.
Published: (2026)
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
by: Niklaus, Joel, et al.
Published: (2026)
by: Niklaus, Joel, et al.
Published: (2026)
The Sustainability Gap in Robotics: A Large-Scale Survey of Sustainability Awareness in 50,000 Research Articles
by: Skuric, Antun, et al.
Published: (2026)
by: Skuric, Antun, et al.
Published: (2026)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
by: Kosson, Atli, et al.
Published: (2024)
by: Kosson, Atli, et al.
Published: (2024)
Towards an empirical understanding of MoE design choices
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
by: Kosson, Atli, et al.
Published: (2023)
by: Kosson, Atli, et al.
Published: (2023)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
by: Fan, Dongyang, et al.
Published: (2024)
by: Fan, Dongyang, et al.
Published: (2024)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
by: Hägele, Alexander, et al.
Published: (2024)
by: Hägele, Alexander, et al.
Published: (2024)
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
by: Allal, Loubna Ben, et al.
Published: (2025)
by: Allal, Loubna Ben, et al.
Published: (2025)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
FineWeb-zhtw: Scalable Curation of Traditional Chinese Text Data from the Web
by: Lin, Cheng-Wei, et al.
Published: (2024)
by: Lin, Cheng-Wei, et al.
Published: (2024)
Latency and Privacy-Aware Resource Allocation in Vehicular Edge Computing
by: Ahmadvand, Hossein, et al.
Published: (2025)
by: Ahmadvand, Hossein, et al.
Published: (2025)
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
by: Wang, Yudong, et al.
Published: (2025)
by: Wang, Yudong, et al.
Published: (2025)
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
by: Matena, Michael, et al.
Published: (2023)
by: Matena, Michael, et al.
Published: (2023)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
by: Deng, Haikang, et al.
Published: (2023)
by: Deng, Haikang, et al.
Published: (2023)
How Do Multilingual Language Models Remember Facts?
by: Fierro, Constanza, et al.
Published: (2024)
by: Fierro, Constanza, et al.
Published: (2024)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
by: Je, Gyung Hyun, et al.
Published: (2025)
by: Je, Gyung Hyun, et al.
Published: (2025)
CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs
by: Nikeghbal, Nafiseh, et al.
Published: (2025)
by: Nikeghbal, Nafiseh, et al.
Published: (2025)
MaskLID: Code-Switching Language Identification through Iterative Masking
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
GIRT-Model: Automated Generation of Issue Report Templates
by: Nikeghbal, Nafiseh, et al.
Published: (2024)
by: Nikeghbal, Nafiseh, et al.
Published: (2024)
GlotScript: A Resource and Tool for Low Resource Writing System Identification
by: Kargaran, Amir Hossein, et al.
Published: (2023)
by: Kargaran, Amir Hossein, et al.
Published: (2023)
Revisiting Multilingual Data Mixtures in Language Model Pretraining
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Soft Merging of Experts with Adaptive Routing
by: Muqeeth, Mohammed, et al.
Published: (2023)
by: Muqeeth, Mohammed, et al.
Published: (2023)
Obsidian : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Obsidian : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Obsidian : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Świeciechów flint : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Obsidian : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Świeciechów flint : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Obsidian : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
by: Bayazit, Deniz, et al.
Published: (2023)
by: Bayazit, Deniz, et al.
Published: (2023)
DABstep: Data Agent Benchmark for Multi-step Reasoning
by: Egg, Alex, et al.
Published: (2025)
by: Egg, Alex, et al.
Published: (2025)
Similar Items
-
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
by: Penedo, Guilherme, et al.
Published: (2024) -
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
by: Messmer, Bettina, et al.
Published: (2025) -
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026) -
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
by: Fan, Dongyang, et al.
Published: (2025) -
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
by: Niklaus, Joel, et al.
Published: (2026)