The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Penedo, Guilherme, Kydlíček, Hynek, allal, Loubna Ben, Lozhkov, Anton, Mitchell, Margaret, Raffel, Colin, Von Werra, Leandro, Wolf, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
by: Penedo, Guilherme, et al.
Published: (2025)
by: Penedo, Guilherme, et al.
Published: (2025)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
by: Niklaus, Joel, et al.
Published: (2026)
by: Niklaus, Joel, et al.
Published: (2026)
FineWeb-zhtw: Scalable Curation of Traditional Chinese Text Data from the Web
by: Lin, Cheng-Wei, et al.
Published: (2024)
by: Lin, Cheng-Wei, et al.
Published: (2024)
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
by: Allal, Loubna Ben, et al.
Published: (2025)
by: Allal, Loubna Ben, et al.
Published: (2025)
The Sustainability Gap in Robotics: A Large-Scale Survey of Sustainability Awareness in 50,000 Research Articles
by: Skuric, Antun, et al.
Published: (2026)
by: Skuric, Antun, et al.
Published: (2026)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
by: Hägele, Alexander, et al.
Published: (2024)
by: Hägele, Alexander, et al.
Published: (2024)
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
by: Wang, Yudong, et al.
Published: (2025)
by: Wang, Yudong, et al.
Published: (2025)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
by: Patel, Ajay, et al.
Published: (2026)
by: Patel, Ajay, et al.
Published: (2026)
FineFreq: A Multilingual Character Frequency Dataset from Web-Scale Text
by: Xu, Binbin
Published: (2025)
by: Xu, Binbin
Published: (2025)
Reward-Augmented Decoding: Efficient Controlled Text Generation With a Unidirectional Reward Model
by: Deng, Haikang, et al.
Published: (2023)
by: Deng, Haikang, et al.
Published: (2023)
Efficiently Estimating Data Efficiency for Language Model Fine-tuning
by: Je, Gyung Hyun, et al.
Published: (2025)
by: Je, Gyung Hyun, et al.
Published: (2025)
SmolVLM: Redefining small and efficient multimodal models
by: Marafioti, Andrés, et al.
Published: (2025)
by: Marafioti, Andrés, et al.
Published: (2025)
The Finest Souls of Our Rivers and Alps
by: Kroll, Paul W.
Published: (2025)
by: Kroll, Paul W.
Published: (2025)
Uncovering Model Processing Strategies with Non-Negative Per-Example Fisher Factorization
by: Matena, Michael, et al.
Published: (2023)
by: Matena, Michael, et al.
Published: (2023)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Fractionation of Edible Fats: Model Fat‐Oil Separation in Decanter Centrifuge
by: Myrofora Kyrimlidou, et al.
Published: (2025)
by: Myrofora Kyrimlidou, et al.
Published: (2025)
Finest positroid subdivisions from maximal weakly separated collections
by: Koshevoy, Gleb A., et al.
Published: (2025)
by: Koshevoy, Gleb A., et al.
Published: (2025)
Poisoning Web-Scale Training Datasets is Practical
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding
by: Bai, Yatong, et al.
Published: (2024)
by: Bai, Yatong, et al.
Published: (2024)
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
by: Fan, Sicheng, et al.
Published: (2026)
by: Fan, Sicheng, et al.
Published: (2026)
Web Element Relocalization in Evolving Web Applications: A Comparative Analysis and Extension Study
by: Kluge, Anton, et al.
Published: (2025)
by: Kluge, Anton, et al.
Published: (2025)
ChineseWebText 2.0: Large-Scale High-quality Chinese Web Text with Multi-dimensional and fine-grained information
by: Zhang, Wanyue, et al.
Published: (2024)
by: Zhang, Wanyue, et al.
Published: (2024)
Towards Best Practices for Open Datasets for LLM Training
by: Baack, Stefan, et al.
Published: (2025)
by: Baack, Stefan, et al.
Published: (2025)
The Organizational Role of Web Services
by: Mitchell, Erik
Published: (2011)
by: Mitchell, Erik
Published: (2011)
Educational Journeys on the Web Frontier.
by: Shneiderman, Ben
Published: (1998)
by: Shneiderman, Ben
Published: (1998)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
Merging by Matching Models in Task Parameter Subspaces
by: Tam, Derek, et al.
Published: (2023)
by: Tam, Derek, et al.
Published: (2023)
Soft Merging of Experts with Adaptive Routing
by: Muqeeth, Mohammed, et al.
Published: (2023)
by: Muqeeth, Mohammed, et al.
Published: (2023)
From Blocking to Breaking: Evaluating the Impact of Adblockers on Web Usability
by: Roongta, Ritik, et al.
Published: (2024)
by: Roongta, Ritik, et al.
Published: (2024)
Refining the Use of the Web (and Web Search) as a Language Teaching and Learning Resource
by: Wu, Shaoqun, et al.
Published: (2009)
by: Wu, Shaoqun, et al.
Published: (2009)
vocalimitationset in WebDataset Format
by: Yadong, Niu
Published: (2025)
by: Yadong, Niu
Published: (2025)
CrediBench: Building Web-Scale Network Datasets for Information Integrity
by: Kondrup, Emma, et al.
Published: (2025)
by: Kondrup, Emma, et al.
Published: (2025)
PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Web Execution Bundles: Reproducible, Accurate, and Archivable Web Measurements
by: Hantke, Florian, et al.
Published: (2025)
by: Hantke, Florian, et al.
Published: (2025)
Measuring Fingerprints of Web-filtered Text Datasets and Fingerprint Propagation Through Training
by: Mansour, Youssef, et al.
Published: (2024)
by: Mansour, Youssef, et al.
Published: (2024)
Obsidian : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Obsidian : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Obsidian : 3D documentation
by: Werra, Dagmara H.
Published: (2026)
by: Werra, Dagmara H.
Published: (2026)
Similar Items
-
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
by: Penedo, Guilherme, et al.
Published: (2025) -
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
by: Niklaus, Joel, et al.
Published: (2026) -
FineWeb-zhtw: Scalable Curation of Traditional Chinese Text Data from the Web
by: Lin, Cheng-Wei, et al.
Published: (2024) -
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
by: Allal, Loubna Ben, et al.
Published: (2025) -
The Sustainability Gap in Robotics: A Large-Scale Survey of Sustainability Awareness in 50,000 Research Articles
by: Skuric, Antun, et al.
Published: (2026)