GneissWeb: Preparing High Quality Data for LLMs at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915417333694464 |
|---|---|
| author | Gohari, Hajar Emami Kadhe, Swanand Ravindra Shah, Syed Yousaf Adam, Constantin Adebayo, Abdulhamid Adusumilli, Praneet Ahmed, Farhan Angel, Nathalie Baracaldo Borse, Santosh Subhashrao Chang, Yuan-Chi Dang, Xuan-Hong Desai, Nirmit Eres, Revital Iwamoto, Ran Karve, Alexei Koyfman, Yan Lee, Wei-Han Liu, Changchang Lublinsky, Boris Ohko, Takuyo Pesce, Pablo Touma, Maroun Wang, Shiqiang Witherspoon, Shalisha Woisetschläger, Herbert Wood, David Wu, Kun-Lung Yoshida, Issei Zawad, Syed Zerfos, Petros Zhou, Yi Bhattacharjee, Bishwaranjan |
| author_facet | Gohari, Hajar Emami Kadhe, Swanand Ravindra Shah, Syed Yousaf Adam, Constantin Adebayo, Abdulhamid Adusumilli, Praneet Ahmed, Farhan Angel, Nathalie Baracaldo Borse, Santosh Subhashrao Chang, Yuan-Chi Dang, Xuan-Hong Desai, Nirmit Eres, Revital Iwamoto, Ran Karve, Alexei Koyfman, Yan Lee, Wei-Han Liu, Changchang Lublinsky, Boris Ohko, Takuyo Pesce, Pablo Touma, Maroun Wang, Shiqiang Witherspoon, Shalisha Woisetschläger, Herbert Wood, David Wu, Kun-Lung Yoshida, Issei Zawad, Syed Zerfos, Petros Zhou, Yi Bhattacharjee, Bishwaranjan |
| contents | Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. Large pre-training datasets for leading LLMs remain inaccessible to the public, whereas many open datasets are small in size (less than 5 trillion tokens), limiting their suitability for training large models.
In this paper, we introduce GneissWeb, a large dataset yielding around 10 trillion tokens that caters to the data quality and quantity requirements of training LLMs. Our GneissWeb recipe that produced the dataset consists of sharded exact sub-string deduplication and a judiciously constructed ensemble of quality filters. GneissWeb achieves a favorable trade-off between data quality and quantity, producing models that outperform models trained on state-of-the-art open large datasets (5+ trillion tokens).
We show that models trained using GneissWeb dataset outperform those trained on FineWeb-V1.1.0 by 2.73 percentage points in terms of average score computed on a set of 11 commonly used benchmarks (both zero-shot and few-shot) for pre-training dataset evaluation. When the evaluation set is extended to 20 benchmarks (both zero-shot and few-shot), models trained using GneissWeb still achieve a 1.75 percentage points advantage over those trained on FineWeb-V1.1.0. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_14907 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | GneissWeb: Preparing High Quality Data for LLMs at Scale Gohari, Hajar Emami Kadhe, Swanand Ravindra Shah, Syed Yousaf Adam, Constantin Adebayo, Abdulhamid Adusumilli, Praneet Ahmed, Farhan Angel, Nathalie Baracaldo Borse, Santosh Subhashrao Chang, Yuan-Chi Dang, Xuan-Hong Desai, Nirmit Eres, Revital Iwamoto, Ran Karve, Alexei Koyfman, Yan Lee, Wei-Han Liu, Changchang Lublinsky, Boris Ohko, Takuyo Pesce, Pablo Touma, Maroun Wang, Shiqiang Witherspoon, Shalisha Woisetschläger, Herbert Wood, David Wu, Kun-Lung Yoshida, Issei Zawad, Syed Zerfos, Petros Zhou, Yi Bhattacharjee, Bishwaranjan Computation and Language Artificial Intelligence Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. Large pre-training datasets for leading LLMs remain inaccessible to the public, whereas many open datasets are small in size (less than 5 trillion tokens), limiting their suitability for training large models. In this paper, we introduce GneissWeb, a large dataset yielding around 10 trillion tokens that caters to the data quality and quantity requirements of training LLMs. Our GneissWeb recipe that produced the dataset consists of sharded exact sub-string deduplication and a judiciously constructed ensemble of quality filters. GneissWeb achieves a favorable trade-off between data quality and quantity, producing models that outperform models trained on state-of-the-art open large datasets (5+ trillion tokens). We show that models trained using GneissWeb dataset outperform those trained on FineWeb-V1.1.0 by 2.73 percentage points in terms of average score computed on a set of 11 commonly used benchmarks (both zero-shot and few-shot) for pre-training dataset evaluation. When the evaluation set is extended to 20 benchmarks (both zero-shot and few-shot), models trained using GneissWeb still achieve a 1.75 percentage points advantage over those trained on FineWeb-V1.1.0. |
| title | GneissWeb: Preparing High Quality Data for LLMs at Scale |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2502.14907 |