Evaluating Mathematical Reasoning Beyond Accuracy
Fuente:
arXiv
Saved in:
| Main Authors: | Xia, Shijie, Li, Xuefeng, Liu, Yixin, Wu, Tongshuang, Liu, Pengfei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LIMO: Less is More for Reasoning
by: Ye, Yixin, et al.
Published: (2025)
by: Ye, Yixin, et al.
Published: (2025)
Progress or Regress? Self-Improvement Reversal in Post-training
by: Wu, Ting, et al.
Published: (2024)
by: Wu, Ting, et al.
Published: (2024)
Beyond Relevance: Evaluate and Improve Retrievers on Perspective Awareness
by: Zhao, Xinran, et al.
Published: (2024)
by: Zhao, Xinran, et al.
Published: (2024)
O1 Replication Journey: A Strategic Progress Report -- Part 1
by: Qin, Yiwei, et al.
Published: (2024)
by: Qin, Yiwei, et al.
Published: (2024)
SAFETY-J: Evaluating Safety with Critique
by: Liu, Yixiu, et al.
Published: (2024)
by: Liu, Yixiu, et al.
Published: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
ToRL: Scaling Tool-Integrated RL
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
Beyond Factual Accuracy: Evaluating Global Reasoning Integrity in RAG Systems with LogicScore
by: Yan, Zhichao, et al.
Published: (2026)
by: Yan, Zhichao, et al.
Published: (2026)
Beyond Rejection Sampling: Trajectory Fusion for Scaling Mathematical Reasoning
by: Deng, Jie, et al.
Published: (2026)
by: Deng, Jie, et al.
Published: (2026)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
OlympicArena Medal Ranks: Who Is the Most Intelligent AI So Far?
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning
by: Yang, Xia, et al.
Published: (2026)
by: Yang, Xia, et al.
Published: (2026)
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
by: Zhou, Zihao, et al.
Published: (2024)
by: Zhou, Zihao, et al.
Published: (2024)
Understanding Reference Policies in Direct Preference Optimization
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
by: Zhang, Jiaqiao, et al.
Published: (2026)
by: Zhang, Jiaqiao, et al.
Published: (2026)
FRoG: Evaluating Fuzzy Reasoning of Generalized Quantifiers in Large Language Models
by: Li, Yiyuan, et al.
Published: (2024)
by: Li, Yiyuan, et al.
Published: (2024)
LIMR: Less is More for RL Scaling
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
by: Huang, Zhen, et al.
Published: (2024)
by: Huang, Zhen, et al.
Published: (2024)
What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code
by: Zhao, Yuze, et al.
Published: (2026)
by: Zhao, Yuze, et al.
Published: (2026)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
Generative AI Act II: Test Time Scaling Drives Cognition Engineering
by: Xia, Shijie, et al.
Published: (2025)
by: Xia, Shijie, et al.
Published: (2025)
DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
by: Wu, Changti, et al.
Published: (2025)
by: Wu, Changti, et al.
Published: (2025)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
by: Wang, Yiming, et al.
Published: (2025)
by: Wang, Yiming, et al.
Published: (2025)
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
by: Liu, Wentao, et al.
Published: (2024)
by: Liu, Wentao, et al.
Published: (2024)
SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking
by: Huang, Weiyang, et al.
Published: (2026)
by: Huang, Weiyang, et al.
Published: (2026)
Beyond Unimodal Shortcuts: MLLMs as Cross-Modal Reasoners for Grounded Named Entity Recognition
by: Ma, Jinlong, et al.
Published: (2026)
by: Ma, Jinlong, et al.
Published: (2026)
One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learning
by: Li, Yiyuan, et al.
Published: (2026)
by: Li, Yiyuan, et al.
Published: (2026)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
by: Lu, Pan, et al.
Published: (2023)
by: Lu, Pan, et al.
Published: (2023)
Toward Automated Robustness Evaluation of Mathematical Reasoning
by: Hou, Yutao, et al.
Published: (2025)
by: Hou, Yutao, et al.
Published: (2025)
Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
by: Zhang, Zhihan, et al.
Published: (2024)
by: Zhang, Zhihan, et al.
Published: (2024)
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
by: Wang, Zengzhi, et al.
Published: (2025)
by: Wang, Zengzhi, et al.
Published: (2025)
THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
by: Chang, Qikai, et al.
Published: (2025)
by: Chang, Qikai, et al.
Published: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
by: Dekoninck, Jasper, et al.
Published: (2026)
by: Dekoninck, Jasper, et al.
Published: (2026)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
by: Zeng, Yixiao, et al.
Published: (2025)
by: Zeng, Yixiao, et al.
Published: (2025)
Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning
by: Zhang, Lan, et al.
Published: (2025)
by: Zhang, Lan, et al.
Published: (2025)
MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
by: Huang, Liangjie, et al.
Published: (2025)
by: Huang, Liangjie, et al.
Published: (2025)
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
by: Patil, Parth, et al.
Published: (2026)
by: Patil, Parth, et al.
Published: (2026)
Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
by: Li, Alan, et al.
Published: (2025)
by: Li, Alan, et al.
Published: (2025)
Similar Items
-
LIMO: Less is More for Reasoning
by: Ye, Yixin, et al.
Published: (2025) -
Progress or Regress? Self-Improvement Reversal in Post-training
by: Wu, Ting, et al.
Published: (2024) -
Beyond Relevance: Evaluate and Improve Retrievers on Perspective Awareness
by: Zhao, Xinran, et al.
Published: (2024) -
O1 Replication Journey: A Strategic Progress Report -- Part 1
by: Qin, Yiwei, et al.
Published: (2024) -
SAFETY-J: Evaluating Safety with Critique
by: Liu, Yixiu, et al.
Published: (2024)