Essential-Web v1.0: 24T tokens of organized web data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: AI, Essential, :, Hojel, Andrew, Pust, Michael, Romanski, Tim, Vanjani, Yash, Kapila, Ritvik, Parmar, Mohit, Chaluvaraju, Adarsh, Tripathy, Alok, Thomas, Anil, Tanwer, Ashish, Shah, Darsh J, Shah, Ishaan, Stratos, Karl, Nguyen, Khoi, Smith, Kurt, Callahan, Michael, Rushton, Peter, Monk, Philip, Mazarakis, Platon, Jamal, Saad, Srivastava, Saurabh, Singla, Somanshu, Vaswani, Ashish
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909653761261568
author AI, Essential
:
Hojel, Andrew
Pust, Michael
Romanski, Tim
Vanjani, Yash
Kapila, Ritvik
Parmar, Mohit
Chaluvaraju, Adarsh
Tripathy, Alok
Thomas, Anil
Tanwer, Ashish
Shah, Darsh J
Shah, Ishaan
Stratos, Karl
Nguyen, Khoi
Smith, Kurt
Callahan, Michael
Rushton, Peter
Monk, Philip
Mazarakis, Platon
Jamal, Saad
Srivastava, Saurabh
Singla, Somanshu
Vaswani, Ashish
author_facet AI, Essential
:
Hojel, Andrew
Pust, Michael
Romanski, Tim
Vanjani, Yash
Kapila, Ritvik
Parmar, Mohit
Chaluvaraju, Adarsh
Tripathy, Alok
Thomas, Anil
Tanwer, Ashish
Shah, Darsh J
Shah, Ishaan
Stratos, Karl
Nguyen, Khoi
Smith, Kurt
Callahan, Michael
Rushton, Peter
Monk, Philip
Mazarakis, Platon
Jamal, Saad
Srivastava, Saurabh
Singla, Somanshu
Vaswani, Ashish
contents Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible data pipelines. We present Essential-Web v1.0, a 24-trillion-token dataset in which every document is annotated with a twelve-category taxonomy covering topic, format, content complexity, and quality. Taxonomy labels are produced by EAI-Distill-0.5b, a fine-tuned 0.5b-parameter model that achieves an annotator agreement within 3% of Qwen2.5-32B-Instruct. With nothing more than SQL-style filters, we obtain competitive web-curated datasets in math (-8.0% relative to SOTA), web code (+14.3%), STEM (+24.5%) and medical (+8.6%). Essential-Web v1.0 is available on HuggingFace: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0
format Preprint
id arxiv_https___arxiv_org_abs_2506_14111
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Essential-Web v1.0: 24T tokens of organized web data
AI, Essential
:
Hojel, Andrew
Pust, Michael
Romanski, Tim
Vanjani, Yash
Kapila, Ritvik
Parmar, Mohit
Chaluvaraju, Adarsh
Tripathy, Alok
Thomas, Anil
Tanwer, Ashish
Shah, Darsh J
Shah, Ishaan
Stratos, Karl
Nguyen, Khoi
Smith, Kurt
Callahan, Michael
Rushton, Peter
Monk, Philip
Mazarakis, Platon
Jamal, Saad
Srivastava, Saurabh
Singla, Somanshu
Vaswani, Ashish
Computation and Language
Artificial Intelligence
Machine Learning
Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible data pipelines. We present Essential-Web v1.0, a 24-trillion-token dataset in which every document is annotated with a twelve-category taxonomy covering topic, format, content complexity, and quality. Taxonomy labels are produced by EAI-Distill-0.5b, a fine-tuned 0.5b-parameter model that achieves an annotator agreement within 3% of Qwen2.5-32B-Instruct. With nothing more than SQL-style filters, we obtain competitive web-curated datasets in math (-8.0% relative to SOTA), web code (+14.3%), STEM (+24.5%) and medical (+8.6%). Essential-Web v1.0 is available on HuggingFace: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0
title Essential-Web v1.0: 24T tokens of organized web data
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.14111