Toward Automated Robustness Evaluation of Mathematical Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hou, Yutao, Xiao, Zeguan, Yu, Fei, Jiang, Yihan, Shuguang, Ma, Dai, Zhaoqian, Huang, Hailiang, Chen, Yun, Chen, Guanhua
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918465408860160
author Hou, Yutao
Xiao, Zeguan
Yu, Fei
Jiang, Yihan
Shuguang, Ma
Dai, Zhaoqian
Huang, Hailiang
Chen, Yun
Chen, Guanhua
author_facet Hou, Yutao
Xiao, Zeguan
Yu, Fei
Jiang, Yihan
Shuguang, Ma
Dai, Zhaoqian
Huang, Hailiang
Chen, Yun
Chen, Guanhua
contents Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks. However, these models exhibit unexpected brittleness, often failing on simple variations of the same underlying task. Existing robustness evaluations predominantly rely on hand-crafted templates or a limited set of perturbation rules. Consequently, such approaches lack the adaptability to probe latent vulnerabilities unique to specific models and remain susceptible to data contamination. To address this, we propose the Math Stress Tester (MaSTer), an automated framework inspired by software stress testing. MaSTer generates adversarial variants via a multi-round rewrite-verify loop, ensuring semantic consistency while successfully inducing model failure. Our framework generates benchmark variants dynamically for each LLM, thus minimizing the risk of data contamination. Experiments on GSM8K and MATH-500 demonstrate the effectiveness of MaSTer on mathematical tasks. Additionally, we validate the framework's extensibility to non-mathematical tasks, highlighting its broad applicability. Furthermore, we demonstrate that the synthesized variants generated by MaSTer can be utilized as a fine-tuning dataset to significantly enhance the model's robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05038
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward Automated Robustness Evaluation of Mathematical Reasoning
Hou, Yutao
Xiao, Zeguan
Yu, Fei
Jiang, Yihan
Shuguang, Ma
Dai, Zhaoqian
Huang, Hailiang
Chen, Yun
Chen, Guanhua
Computation and Language
Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks. However, these models exhibit unexpected brittleness, often failing on simple variations of the same underlying task. Existing robustness evaluations predominantly rely on hand-crafted templates or a limited set of perturbation rules. Consequently, such approaches lack the adaptability to probe latent vulnerabilities unique to specific models and remain susceptible to data contamination. To address this, we propose the Math Stress Tester (MaSTer), an automated framework inspired by software stress testing. MaSTer generates adversarial variants via a multi-round rewrite-verify loop, ensuring semantic consistency while successfully inducing model failure. Our framework generates benchmark variants dynamically for each LLM, thus minimizing the risk of data contamination. Experiments on GSM8K and MATH-500 demonstrate the effectiveness of MaSTer on mathematical tasks. Additionally, we validate the framework's extensibility to non-mathematical tasks, highlighting its broad applicability. Furthermore, we demonstrate that the synthesized variants generated by MaSTer can be utilized as a fine-tuning dataset to significantly enhance the model's robustness.
title Toward Automated Robustness Evaluation of Mathematical Reasoning
topic Computation and Language
url https://arxiv.org/abs/2506.05038