Do LLMs Overthink Basic Math Reasoning? Benchmarking the Accuracy-Efficiency Tradeoff in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Srivastava, Gaurav, Hussain, Aafiya, Srinivasan, Sriram, Wang, Xuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EffGen: Enabling Small Language Models as Capable Autonomous Agents
by: Srivastava, Gaurav, et al.
Published: (2026)
by: Srivastava, Gaurav, et al.
Published: (2026)
BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models
by: Srivastava, Gaurav, et al.
Published: (2025)
by: Srivastava, Gaurav, et al.
Published: (2025)
Batch Prompting Suppresses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models
by: Srivastava, Saurabh, et al.
Published: (2025)
by: Srivastava, Saurabh, et al.
Published: (2025)
SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
by: Hussain, Aafiya, et al.
Published: (2026)
by: Hussain, Aafiya, et al.
Published: (2026)
Towards Reasoning Ability of Small Language Models
by: Srivastava, Gaurav, et al.
Published: (2025)
by: Srivastava, Gaurav, et al.
Published: (2025)
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
by: Pu, Xiao, et al.
Published: (2025)
by: Pu, Xiao, et al.
Published: (2025)
Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis
by: Yadav, Anushka, et al.
Published: (2025)
by: Yadav, Anushka, et al.
Published: (2025)
JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
by: Bi, Zhenyu, et al.
Published: (2025)
by: Bi, Zhenyu, et al.
Published: (2025)
DEBATE, TRAIN, EVOLVE: Self Evolution of Language Model Reasoning
by: Srivastava, Gaurav, et al.
Published: (2025)
by: Srivastava, Gaurav, et al.
Published: (2025)
DRQA: Dynamic Reasoning Quota Allocation for Controlling Overthinking in Reasoning Large Language Models
by: Yan, Kaiwen, et al.
Published: (2025)
by: Yan, Kaiwen, et al.
Published: (2025)
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
by: Sui, Yang, et al.
Published: (2025)
by: Sui, Yang, et al.
Published: (2025)
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
by: Xue, Boyang, et al.
Published: (2025)
by: Xue, Boyang, et al.
Published: (2025)
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
by: Chen, Xingyu, et al.
Published: (2024)
by: Chen, Xingyu, et al.
Published: (2024)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
by: Wang, Lei, et al.
Published: (2024)
by: Wang, Lei, et al.
Published: (2024)
Benchmarking Large Language Models for Math Reasoning Tasks
by: Seßler, Kathrin, et al.
Published: (2024)
by: Seßler, Kathrin, et al.
Published: (2024)
SLIM-LLMs: Modeling of Style-Sensory Language RelationshipsThrough Low-Dimensional Representations
by: Khalid, Osama, et al.
Published: (2025)
by: Khalid, Osama, et al.
Published: (2025)
Think, But Don't Overthink: Reproducing Recursive Language Models
by: Wang, Daren
Published: (2026)
by: Wang, Daren
Published: (2026)
Mitigating Overthinking in Large Reasoning Language Models via Reasoning Path Deviation Monitoring
by: Guan, Weixin, et al.
Published: (2026)
by: Guan, Weixin, et al.
Published: (2026)
Mitigating Overthinking through Reasoning Shaping
by: Song, Feifan, et al.
Published: (2025)
by: Song, Feifan, et al.
Published: (2025)
Reasoning or Overthinking: Evaluating Large Language Models on Financial Sentiment Analysis
by: Vamvourellis, Dimitris, et al.
Published: (2025)
by: Vamvourellis, Dimitris, et al.
Published: (2025)
Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
by: Xu, Liang, et al.
Published: (2024)
by: Xu, Liang, et al.
Published: (2024)
Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
by: Sun, Haoxiang, et al.
Published: (2025)
by: Sun, Haoxiang, et al.
Published: (2025)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
by: Ying, Huaiyuan, et al.
Published: (2024)
by: Ying, Huaiyuan, et al.
Published: (2024)
BadReasoner: Planting Tunable Overthinking Backdoors into Large Reasoning Models for Fun or Profit
by: Yi, Biao, et al.
Published: (2025)
by: Yi, Biao, et al.
Published: (2025)
Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization
by: Diao, Xingjian, et al.
Published: (2026)
by: Diao, Xingjian, et al.
Published: (2026)
Avoiding Overthinking and Underthinking: Curriculum-Aware Budget Scheduling for LLMs
by: Rahman, Amirul, et al.
Published: (2026)
by: Rahman, Amirul, et al.
Published: (2026)
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
by: Zhang, Xinliang Frederick, et al.
Published: (2025)
by: Zhang, Xinliang Frederick, et al.
Published: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
by: Huang, Kaixuan, et al.
Published: (2025)
by: Huang, Kaixuan, et al.
Published: (2025)
SafeMath: Inference-time Safety improves Math Accuracy
by: Basu, Sagnik, et al.
Published: (2026)
by: Basu, Sagnik, et al.
Published: (2026)
Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking
by: Han, Jinyi, et al.
Published: (2025)
by: Han, Jinyi, et al.
Published: (2025)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
by: Zhao, Yilun, et al.
Published: (2023)
by: Zhao, Yilun, et al.
Published: (2023)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
by: Tian, Shi-Yu, et al.
Published: (2025)
by: Tian, Shi-Yu, et al.
Published: (2025)
Mitigating Overthinking in Large Reasoning Models via Manifold Steering
by: Huang, Yao, et al.
Published: (2025)
by: Huang, Yao, et al.
Published: (2025)
Simpson's Paradox and the Accuracy-Fluency Tradeoff in Translation
by: Lim, Zheng Wei, et al.
Published: (2024)
by: Lim, Zheng Wei, et al.
Published: (2024)
Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learning
by: Wang, Qianyue, et al.
Published: (2026)
by: Wang, Qianyue, et al.
Published: (2026)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026)
by: Dekoninck, Jasper, et al.
Published: (2026)
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
by: Xu, Haoran, et al.
Published: (2025)
by: Xu, Haoran, et al.
Published: (2025)
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
by: Guan, Xinyu, et al.
Published: (2025)
by: Guan, Xinyu, et al.
Published: (2025)
STEM-POM: Evaluating Language Models Math-Symbol Reasoning in Document Parsing
by: Zou, Jiaru, et al.
Published: (2024)
by: Zou, Jiaru, et al.
Published: (2024)
Similar Items
-
EffGen: Enabling Small Language Models as Capable Autonomous Agents
by: Srivastava, Gaurav, et al.
Published: (2026) -
BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models
by: Srivastava, Gaurav, et al.
Published: (2025) -
Batch Prompting Suppresses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models
by: Srivastava, Saurabh, et al.
Published: (2025) -
SoundBreak: A Systematic Study of Audio-Only Adversarial Attacks on Trimodal Models
by: Hussain, Aafiya, et al.
Published: (2026) -
Towards Reasoning Ability of Small Language Models
by: Srivastava, Gaurav, et al.
Published: (2025)