MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Hongwei, Zheng, Zilong, Qiao, Yuxuan, Duan, Haodong, Fei, Zhiwei, Zhou, Fengzhe, Zhang, Wenwei, Zhang, Songyang, Lin, Dahua, Chen, Kai |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
par: Wang, Chonghua, et autres
Publié: (2024)
par: Wang, Chonghua, et autres
Publié: (2024)
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
par: Zhang, Chuyu, et autres
Publié: (2024)
par: Zhang, Chuyu, et autres
Publié: (2024)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
par: Zhuo, Jingming, et autres
Publié: (2024)
par: Zhuo, Jingming, et autres
Publié: (2024)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
par: Li, Mo, et autres
Publié: (2024)
par: Li, Mo, et autres
Publié: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
par: Ying, Huaiyuan, et autres
Publié: (2024)
par: Ying, Huaiyuan, et autres
Publié: (2024)
InternLM-Law: An Open Source Chinese Legal Large Language Model
par: Fei, Zhiwei, et autres
Publié: (2024)
par: Fei, Zhiwei, et autres
Publié: (2024)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
par: Cao, Maosong, et autres
Publié: (2024)
par: Cao, Maosong, et autres
Publié: (2024)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
par: Qiao, Yuxuan, et autres
Publié: (2024)
par: Qiao, Yuxuan, et autres
Publié: (2024)
Are Your LLMs Capable of Stable Reasoning?
par: Liu, Junnan, et autres
Publié: (2024)
par: Liu, Junnan, et autres
Publié: (2024)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
par: Liu, Shudong, et autres
Publié: (2025)
par: Liu, Shudong, et autres
Publié: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
par: Li, Xin, et autres
Publié: (2025)
par: Li, Xin, et autres
Publié: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
par: Dekoninck, Jasper, et autres
Publié: (2026)
par: Dekoninck, Jasper, et autres
Publié: (2026)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
par: Lyu, Chengqi, et autres
Publié: (2025)
par: Lyu, Chengqi, et autres
Publié: (2025)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
par: Fang, Meng, et autres
Publié: (2024)
par: Fang, Meng, et autres
Publié: (2024)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
par: Gu, Yuzhe, et autres
Publié: (2025)
par: Gu, Yuzhe, et autres
Publié: (2025)
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
par: Zhou, Baichuan, et autres
Publié: (2024)
par: Zhou, Baichuan, et autres
Publié: (2024)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
par: Wang, Yiming, et autres
Publié: (2025)
par: Wang, Yiming, et autres
Publié: (2025)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
par: Li, Lijun, et autres
Publié: (2024)
par: Li, Lijun, et autres
Publié: (2024)
PAI-Bench: A Comprehensive Benchmark For Physical AI
par: Zhou, Fengzhe, et autres
Publié: (2025)
par: Zhou, Fengzhe, et autres
Publié: (2025)
JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)
par: Amin, Nishil, et autres
Publié: (2026)
par: Amin, Nishil, et autres
Publié: (2026)
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
par: Hua, Zhouqi, et autres
Publié: (2025)
par: Hua, Zhouqi, et autres
Publié: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
par: Ding, Shuangrui, et autres
Publié: (2026)
par: Ding, Shuangrui, et autres
Publié: (2026)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
par: Chen, Zehui, et autres
Publié: (2023)
par: Chen, Zehui, et autres
Publié: (2023)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
par: Wang, Lei, et autres
Publié: (2024)
par: Wang, Lei, et autres
Publié: (2024)
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
par: Zhan, Shaoxiong, et autres
Publié: (2025)
par: Zhan, Shaoxiong, et autres
Publié: (2025)
RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
par: Zhang, Jie, et autres
Publié: (2025)
par: Zhang, Jie, et autres
Publié: (2025)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
par: Li, Rongjie, et autres
Publié: (2024)
par: Li, Rongjie, et autres
Publié: (2024)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
par: Fang, Xinyu, et autres
Publié: (2024)
par: Fang, Xinyu, et autres
Publié: (2024)
Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark
par: Mercorio, Fabio, et autres
Publié: (2024)
par: Mercorio, Fabio, et autres
Publié: (2024)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
par: Zou, Chengke, et autres
Publié: (2024)
par: Zou, Chengke, et autres
Publié: (2024)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
par: Li, Can, et autres
Publié: (2025)
par: Li, Can, et autres
Publié: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
par: Li, Wei, et autres
Publié: (2024)
par: Li, Wei, et autres
Publié: (2024)
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
par: Yang, Tengchao, et autres
Publié: (2025)
par: Yang, Tengchao, et autres
Publié: (2025)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
par: Zhang, Yuanhe, et autres
Publié: (2025)
par: Zhang, Yuanhe, et autres
Publié: (2025)
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
par: Glazer, Elliot, et autres
Publié: (2024)
par: Glazer, Elliot, et autres
Publié: (2024)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
par: Tong, Yuxuan, et autres
Publié: (2024)
par: Tong, Yuxuan, et autres
Publié: (2024)
Rectifying LLM Thought from Lens of Optimization
par: Liu, Junnan, et autres
Publié: (2025)
par: Liu, Junnan, et autres
Publié: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
par: Li, Xiaoyuan, et autres
Publié: (2025)
par: Li, Xiaoyuan, et autres
Publié: (2025)
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
par: Yang, Sihan, et autres
Publié: (2025)
par: Yang, Sihan, et autres
Publié: (2025)
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
par: Zhao, Haiteng, et autres
Publié: (2025)
par: Zhao, Haiteng, et autres
Publié: (2025)
Documents similaires
-
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
par: Wang, Chonghua, et autres
Publié: (2024) -
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
par: Zhang, Chuyu, et autres
Publié: (2024) -
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
par: Zhuo, Jingming, et autres
Publié: (2024) -
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
par: Li, Mo, et autres
Publié: (2024) -
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
par: Ying, Huaiyuan, et autres
Publié: (2024)