MathClean: A Benchmark for Synthetic Mathematical Data Cleaning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Hao, Qiang, Meiyi, Li, Yuying, He, Zefeng, Guo, Yongzhen, Zhu, Zhengzhou, Zhang, Wentao, Cui, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909511997980672
author Liang, Hao
Qiang, Meiyi
Li, Yuying
He, Zefeng
Guo, Yongzhen
Zhu, Zhengzhou
Zhang, Wentao
Cui, Bin
author_facet Liang, Hao
Qiang, Meiyi
Li, Yuying
He, Zefeng
Guo, Yongzhen
Zhu, Zhengzhou
Zhang, Wentao
Cui, Bin
contents With the rapid development of large language models (LLMs), the quality of training data has become crucial. Among the various types of training data, mathematical data plays a key role in enabling LLMs to acquire strong reasoning abilities. While high-quality open-source data is important, it is often insufficient for pre-training, necessitating the addition of synthetic math problems. However, synthetic math questions and answers can introduce inaccuracies, which may degrade both the training data and web data. Therefore, an effective method for cleaning synthetic math data is essential. In this paper, we propose the MathClean benchmark to evaluate the effectiveness of math data cleaning models. The MathClean benchmark consists of 2,000 correct questions and 2,000 erroneous questions with additional 2,000 correct and erroneous answers sourced from augmented data based on GSM8K and MATH. Moreover, we also annotate error types for each question or answer, since it can assess whether models can correctly identify the error categories for future improvements. Finally, we present comprehensive evaluations using state-of-the-art (SOTA) models. Our results demonstrate that even strong models like GPT-o1 and DeepSeek-R1 perform poorly on this benchmark, highlighting the utility of MathClean. Our code and data is available at https://github.com/YuYingLi0/MathClean.
format Preprint
id arxiv_https___arxiv_org_abs_2502_19058
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
Liang, Hao
Qiang, Meiyi
Li, Yuying
He, Zefeng
Guo, Yongzhen
Zhu, Zhengzhou
Zhang, Wentao
Cui, Bin
Computation and Language
With the rapid development of large language models (LLMs), the quality of training data has become crucial. Among the various types of training data, mathematical data plays a key role in enabling LLMs to acquire strong reasoning abilities. While high-quality open-source data is important, it is often insufficient for pre-training, necessitating the addition of synthetic math problems. However, synthetic math questions and answers can introduce inaccuracies, which may degrade both the training data and web data. Therefore, an effective method for cleaning synthetic math data is essential. In this paper, we propose the MathClean benchmark to evaluate the effectiveness of math data cleaning models. The MathClean benchmark consists of 2,000 correct questions and 2,000 erroneous questions with additional 2,000 correct and erroneous answers sourced from augmented data based on GSM8K and MATH. Moreover, we also annotate error types for each question or answer, since it can assess whether models can correctly identify the error categories for future improvements. Finally, we present comprehensive evaluations using state-of-the-art (SOTA) models. Our results demonstrate that even strong models like GPT-o1 and DeepSeek-R1 perform poorly on this benchmark, highlighting the utility of MathClean. Our code and data is available at https://github.com/YuYingLi0/MathClean.
title MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
topic Computation and Language
url https://arxiv.org/abs/2502.19058