DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Zhiwei, Liang, Tian, Xu, Jiahao, Liu, Qiuzhi, Chen, Xingyu, Wang, Yue, Song, Linfeng, Yu, Dian, Liang, Zhenwen, Wang, Wenxuan, Zhang, Zhuosheng, Wang, Rui, Tu, Zhaopeng, Mi, Haitao, Yu, Dong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913854352523264
author He, Zhiwei
Liang, Tian
Xu, Jiahao
Liu, Qiuzhi
Chen, Xingyu
Wang, Yue
Song, Linfeng
Yu, Dian
Liang, Zhenwen
Wang, Wenxuan
Zhang, Zhuosheng
Wang, Rui
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
author_facet He, Zhiwei
Liang, Tian
Xu, Jiahao
Liu, Qiuzhi
Chen, Xingyu
Wang, Yue
Song, Linfeng
Yu, Dian
Liang, Zhenwen
Wang, Wenxuan
Zhang, Zhuosheng
Wang, Rui
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
contents Reinforcement learning (RL) with large language models shows promise in complex reasoning. However, its progress is hindered by the lack of large-scale training data that is sufficiently challenging, contamination-free and verifiable. To this end, we introduce DeepMath-103K, a large-scale mathematical dataset designed with high difficulty (primarily levels 5-9), rigorous decontamination against numerous benchmarks, and verifiable answers for rule-based RL reward. It further includes three distinct R1 solutions adaptable for diverse training paradigms such as supervised fine-tuning (SFT). Spanning a wide range of mathematical topics, DeepMath-103K fosters the development of generalizable and advancing reasoning. Notably, models trained on DeepMath-103K achieve state-of-the-art results on challenging mathematical benchmarks and demonstrate generalization beyond math such as biology, physics and chemistry, underscoring its broad efficacy. Data: https://huggingface.co/datasets/zwhe99/DeepMath-103K.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
He, Zhiwei
Liang, Tian
Xu, Jiahao
Liu, Qiuzhi
Chen, Xingyu
Wang, Yue
Song, Linfeng
Yu, Dian
Liang, Zhenwen
Wang, Wenxuan
Zhang, Zhuosheng
Wang, Rui
Tu, Zhaopeng
Mi, Haitao
Yu, Dong
Computation and Language
Artificial Intelligence
Reinforcement learning (RL) with large language models shows promise in complex reasoning. However, its progress is hindered by the lack of large-scale training data that is sufficiently challenging, contamination-free and verifiable. To this end, we introduce DeepMath-103K, a large-scale mathematical dataset designed with high difficulty (primarily levels 5-9), rigorous decontamination against numerous benchmarks, and verifiable answers for rule-based RL reward. It further includes three distinct R1 solutions adaptable for diverse training paradigms such as supervised fine-tuning (SFT). Spanning a wide range of mathematical topics, DeepMath-103K fosters the development of generalizable and advancing reasoning. Notably, models trained on DeepMath-103K achieve state-of-the-art results on challenging mathematical benchmarks and demonstrate generalization beyond math such as biology, physics and chemistry, underscoring its broad efficacy. Data: https://huggingface.co/datasets/zwhe99/DeepMath-103K.
title DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.11456