SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Liang, Xue, Hang, Zhu, Lei, Zhao, Kangkang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910315753504768
author Xu, Liang
Xue, Hang
Zhu, Lei
Zhao, Kangkang
author_facet Xu, Liang
Xue, Hang
Zhu, Lei
Zhao, Kangkang
contents We introduce SuperCLUE-Math6(SC-Math6), a new benchmark dataset to evaluate the mathematical reasoning abilities of Chinese language models. SC-Math6 is designed as an upgraded Chinese version of the GSM8K dataset with enhanced difficulty, diversity, and application scope. It consists of over 2000 mathematical word problems requiring multi-step reasoning and providing natural language solutions. We propose an innovative scheme to quantify the reasoning capability of large models based on performance over problems with different reasoning steps. Experiments on 13 representative Chinese models demonstrate a clear stratification of reasoning levels, with top models like GPT-4 showing superior performance. SC-Math6 fills the gap in Chinese mathematical reasoning benchmarks and provides a comprehensive testbed to advance the intelligence of Chinese language models.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11819
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
Xu, Liang
Xue, Hang
Zhu, Lei
Zhao, Kangkang
Computation and Language
Artificial Intelligence
We introduce SuperCLUE-Math6(SC-Math6), a new benchmark dataset to evaluate the mathematical reasoning abilities of Chinese language models. SC-Math6 is designed as an upgraded Chinese version of the GSM8K dataset with enhanced difficulty, diversity, and application scope. It consists of over 2000 mathematical word problems requiring multi-step reasoning and providing natural language solutions. We propose an innovative scheme to quantify the reasoning capability of large models based on performance over problems with different reasoning steps. Experiments on 13 representative Chinese models demonstrate a clear stratification of reasoning levels, with top models like GPT-4 showing superior performance. SC-Math6 fills the gap in Chinese mathematical reasoning benchmarks and provides a comprehensive testbed to advance the intelligence of Chinese language models.
title SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.11819