An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hao, Yuren, Wan, Xiang, Zhai, ChengXiang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
par: Pei, Qizhi, et autres
Publié: (2025)
par: Pei, Qizhi, et autres
Publié: (2025)
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction
par: Hao, Yuren, et autres
Publié: (2026)
par: Hao, Yuren, et autres
Publié: (2026)
Evaluating Robustness of Reward Models for Mathematical Reasoning
par: Kim, Sunghwan, et autres
Publié: (2024)
par: Kim, Sunghwan, et autres
Publié: (2024)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
par: Lai, Yuhang, et autres
Publié: (2026)
par: Lai, Yuhang, et autres
Publié: (2026)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
par: Li, Ming, et autres
Publié: (2025)
par: Li, Ming, et autres
Publié: (2025)
ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges
par: Qian, Cheng, et autres
Publié: (2025)
par: Qian, Cheng, et autres
Publié: (2025)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
par: Li, Chengpeng, et autres
Publié: (2024)
par: Li, Chengpeng, et autres
Publié: (2024)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
par: Leang, Joshua Ong Jun, et autres
Publié: (2024)
par: Leang, Joshua Ong Jun, et autres
Publié: (2024)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
par: Yu, Erxin, et autres
Publié: (2025)
par: Yu, Erxin, et autres
Publié: (2025)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
par: Davoodi, Arash Gholami, et autres
Publié: (2024)
par: Davoodi, Arash Gholami, et autres
Publié: (2024)
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
par: Singh, Joykirat, et autres
Publié: (2024)
par: Singh, Joykirat, et autres
Publié: (2024)
PromptCoT: Synthesizing Olympiad-level Problems for Mathematical Reasoning in Large Language Models
par: Zhao, Xueliang, et autres
Publié: (2025)
par: Zhao, Xueliang, et autres
Publié: (2025)
Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoning
par: Chegini, Atoosa, et autres
Publié: (2026)
par: Chegini, Atoosa, et autres
Publié: (2026)
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
par: Huang, Yuzhen, et autres
Publié: (2025)
par: Huang, Yuzhen, et autres
Publié: (2025)
PrAg-PO: Prompt Augmented Policy Optimization for Robust and Diverse Mathematical Reasoning
par: Lu, Wenquan, et autres
Publié: (2026)
par: Lu, Wenquan, et autres
Publié: (2026)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
par: Zhang, Zhenru, et autres
Publié: (2025)
par: Zhang, Zhenru, et autres
Publié: (2025)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
par: Zheng, Chujie, et autres
Publié: (2024)
par: Zheng, Chujie, et autres
Publié: (2024)
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
par: Wang, Yiming, et autres
Publié: (2024)
par: Wang, Yiming, et autres
Publié: (2024)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
par: Tang, Zhengyang, et autres
Publié: (2024)
par: Tang, Zhengyang, et autres
Publié: (2024)
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs
par: Li, Yinghui, et autres
Publié: (2025)
par: Li, Yinghui, et autres
Publié: (2025)
CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
par: Mao, Yujun, et autres
Publié: (2024)
par: Mao, Yujun, et autres
Publié: (2024)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
par: Ma, Jingkun, et autres
Publié: (2024)
par: Ma, Jingkun, et autres
Publié: (2024)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
par: Guo, Siyuan, et autres
Publié: (2024)
par: Guo, Siyuan, et autres
Publié: (2024)
Can A Gamer Train A Mathematical Reasoning Model?
par: Shin, Andrew
Publié: (2025)
par: Shin, Andrew
Publié: (2025)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
par: Luo, Ruilin, et autres
Publié: (2025)
par: Luo, Ruilin, et autres
Publié: (2025)
Integrating Arithmetic Learning Improves Mathematical Reasoning in Smaller Models
par: Gangwar, Neeraj, et autres
Publié: (2025)
par: Gangwar, Neeraj, et autres
Publié: (2025)
Stepwise Self-Consistent Mathematical Reasoning with Large Language Models
par: Zhao, Zilong, et autres
Publié: (2024)
par: Zhao, Zilong, et autres
Publié: (2024)
Forward-Backward Reasoning in Large Language Models for Mathematical Verification
par: Jiang, Weisen, et autres
Publié: (2023)
par: Jiang, Weisen, et autres
Publié: (2023)
Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
par: Noël, Valentin
Publié: (2026)
par: Noël, Valentin
Publié: (2026)
Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
par: Yu, Fengming, et autres
Publié: (2025)
par: Yu, Fengming, et autres
Publié: (2025)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
par: Zeng, Liang, et autres
Publié: (2024)
par: Zeng, Liang, et autres
Publié: (2024)
Mathematical Reasoning via Intervention-Based Time-Series Causal Discovery Using LLMs as Concept Mastery Simulators
par: Okita, Tsuyoshi
Publié: (2026)
par: Okita, Tsuyoshi
Publié: (2026)
Towards Robust Mathematical Reasoning
par: Luong, Thang, et autres
Publié: (2025)
par: Luong, Thang, et autres
Publié: (2025)
Towards Autonomous Mathematics Research
par: Feng, Tony, et autres
Publié: (2026)
par: Feng, Tony, et autres
Publié: (2026)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
par: Zheng, Xiang, et autres
Publié: (2026)
par: Zheng, Xiang, et autres
Publié: (2026)
PREMISE: Scalable and Strategic Prompt Optimization for Efficient Mathematical Reasoning in Large Models
par: Yu, Ye, et autres
Publié: (2025)
par: Yu, Ye, et autres
Publié: (2025)
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models
par: Singh, Joykirat, et autres
Publié: (2025)
par: Singh, Joykirat, et autres
Publié: (2025)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
par: Lu, Pan, et autres
Publié: (2023)
par: Lu, Pan, et autres
Publié: (2023)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
par: Shao, Zhihong, et autres
Publié: (2024)
par: Shao, Zhihong, et autres
Publié: (2024)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
par: Dipta, Shubhashis Roy, et autres
Publié: (2026)
par: Dipta, Shubhashis Roy, et autres
Publié: (2026)
Documents similaires
-
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
par: Pei, Qizhi, et autres
Publié: (2025) -
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction
par: Hao, Yuren, et autres
Publié: (2026) -
Evaluating Robustness of Reward Models for Mathematical Reasoning
par: Kim, Sunghwan, et autres
Publié: (2024) -
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
par: Lai, Yuhang, et autres
Publié: (2026) -
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
par: Li, Ming, et autres
Publié: (2025)