MathPile: A Billion-Token-Scale Pretraining Corpus for Math

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Zengzhi, Li, Xuefeng, Xia, Rui, Liu, Pengfei
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916458385113088
author Wang, Zengzhi
Li, Xuefeng
Xia, Rui
Liu, Pengfei
author_facet Wang, Zengzhi
Li, Xuefeng
Xia, Rui
Liu, Pengfei
contents High-quality, large-scale corpora are the cornerstone of building foundation models. In this work, we introduce MathPile, a diverse and high-quality math-centric corpus comprising about 9.5 billion tokens. Throughout its creation, we adhered to the principle of "less is more", firmly believing in the supremacy of data quality over quantity, even in the pre-training phase. Our meticulous data collection and processing efforts included a complex suite of preprocessing, prefiltering, language identification, cleaning, filtering, and deduplication, ensuring the high quality of our corpus. Furthermore, we performed data contamination detection on downstream benchmark test sets to eliminate duplicates and conducted continual pre-training experiments, booting the performance on common mathematical reasoning benchmarks. We aim for our MathPile to boost language models' mathematical reasoning abilities and open-source its different versions and processing scripts to advance the field.
format Preprint
id arxiv_https___arxiv_org_abs_2312_17120
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MathPile: A Billion-Token-Scale Pretraining Corpus for Math
Wang, Zengzhi
Li, Xuefeng
Xia, Rui
Liu, Pengfei
Computation and Language
Artificial Intelligence
Machine Learning
High-quality, large-scale corpora are the cornerstone of building foundation models. In this work, we introduce MathPile, a diverse and high-quality math-centric corpus comprising about 9.5 billion tokens. Throughout its creation, we adhered to the principle of "less is more", firmly believing in the supremacy of data quality over quantity, even in the pre-training phase. Our meticulous data collection and processing efforts included a complex suite of preprocessing, prefiltering, language identification, cleaning, filtering, and deduplication, ensuring the high quality of our corpus. Furthermore, we performed data contamination detection on downstream benchmark test sets to eliminate duplicates and conducted continual pre-training experiments, booting the performance on common mathematical reasoning benchmarks. We aim for our MathPile to boost language models' mathematical reasoning abilities and open-source its different versions and processing scripts to advance the field.
title MathPile: A Billion-Token-Scale Pretraining Corpus for Math
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.17120