MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Hongwei, Zheng, Zilong, Qiao, Yuxuan, Duan, Haodong, Fei, Zhiwei, Zhou, Fengzhe, Zhang, Wenwei, Zhang, Songyang, Lin, Dahua, Chen, Kai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
von: Wang, Chonghua, et al.
Veröffentlicht: (2024)
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
von: Zhang, Chuyu, et al.
Veröffentlicht: (2024)
von: Zhang, Chuyu, et al.
Veröffentlicht: (2024)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024)
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
von: Li, Mo, et al.
Veröffentlicht: (2024)
von: Li, Mo, et al.
Veröffentlicht: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
InternLM-Law: An Open Source Chinese Legal Large Language Model
von: Fei, Zhiwei, et al.
Veröffentlicht: (2024)
von: Fei, Zhiwei, et al.
Veröffentlicht: (2024)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
von: Cao, Maosong, et al.
Veröffentlicht: (2024)
von: Cao, Maosong, et al.
Veröffentlicht: (2024)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2024)
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2024)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
von: Liu, Shudong, et al.
Veröffentlicht: (2025)
von: Liu, Shudong, et al.
Veröffentlicht: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
von: Lyu, Chengqi, et al.
Veröffentlicht: (2025)
von: Lyu, Chengqi, et al.
Veröffentlicht: (2025)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
von: Fang, Meng, et al.
Veröffentlicht: (2024)
von: Fang, Meng, et al.
Veröffentlicht: (2024)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
von: Gu, Yuzhe, et al.
Veröffentlicht: (2025)
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
von: Zhou, Baichuan, et al.
Veröffentlicht: (2024)
von: Zhou, Baichuan, et al.
Veröffentlicht: (2024)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
von: Li, Lijun, et al.
Veröffentlicht: (2024)
von: Li, Lijun, et al.
Veröffentlicht: (2024)
PAI-Bench: A Comprehensive Benchmark For Physical AI
von: Zhou, Fengzhe, et al.
Veröffentlicht: (2025)
von: Zhou, Fengzhe, et al.
Veröffentlicht: (2025)
JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)
von: Amin, Nishil, et al.
Veröffentlicht: (2026)
von: Amin, Nishil, et al.
Veröffentlicht: (2026)
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
von: Hua, Zhouqi, et al.
Veröffentlicht: (2025)
von: Hua, Zhouqi, et al.
Veröffentlicht: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
von: Ding, Shuangrui, et al.
Veröffentlicht: (2026)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2025)
von: Zhan, Shaoxiong, et al.
Veröffentlicht: (2025)
RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
von: Li, Rongjie, et al.
Veröffentlicht: (2024)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark
von: Mercorio, Fabio, et al.
Veröffentlicht: (2024)
von: Mercorio, Fabio, et al.
Veröffentlicht: (2024)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
von: Li, Can, et al.
Veröffentlicht: (2025)
von: Li, Can, et al.
Veröffentlicht: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
von: Yang, Tengchao, et al.
Veröffentlicht: (2025)
von: Yang, Tengchao, et al.
Veröffentlicht: (2025)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2025)
von: Zhang, Yuanhe, et al.
Veröffentlicht: (2025)
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
von: Glazer, Elliot, et al.
Veröffentlicht: (2024)
von: Glazer, Elliot, et al.
Veröffentlicht: (2024)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
von: Tong, Yuxuan, et al.
Veröffentlicht: (2024)
von: Tong, Yuxuan, et al.
Veröffentlicht: (2024)
Rectifying LLM Thought from Lens of Optimization
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
von: Liu, Junnan, et al.
Veröffentlicht: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
von: Yang, Sihan, et al.
Veröffentlicht: (2025)
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
von: Zhao, Haiteng, et al.
Veröffentlicht: (2025)
von: Zhao, Haiteng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
von: Wang, Chonghua, et al.
Veröffentlicht: (2024) -
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
von: Zhang, Chuyu, et al.
Veröffentlicht: (2024) -
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024) -
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
von: Li, Mo, et al.
Veröffentlicht: (2024) -
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)