MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Hongwei, Zheng, Zilong, Qiao, Yuxuan, Duan, Haodong, Fei, Zhiwei, Zhou, Fengzhe, Zhang, Wenwei, Zhang, Songyang, Lin, Dahua, Chen, Kai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
di: Wang, Chonghua, et al.
Pubblicazione: (2024)
di: Wang, Chonghua, et al.
Pubblicazione: (2024)
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
di: Zhang, Chuyu, et al.
Pubblicazione: (2024)
di: Zhang, Chuyu, et al.
Pubblicazione: (2024)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
di: Zhuo, Jingming, et al.
Pubblicazione: (2024)
di: Zhuo, Jingming, et al.
Pubblicazione: (2024)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
di: Li, Mo, et al.
Pubblicazione: (2024)
di: Li, Mo, et al.
Pubblicazione: (2024)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
di: Ying, Huaiyuan, et al.
Pubblicazione: (2024)
di: Ying, Huaiyuan, et al.
Pubblicazione: (2024)
InternLM-Law: An Open Source Chinese Legal Large Language Model
di: Fei, Zhiwei, et al.
Pubblicazione: (2024)
di: Fei, Zhiwei, et al.
Pubblicazione: (2024)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
di: Cao, Maosong, et al.
Pubblicazione: (2024)
di: Cao, Maosong, et al.
Pubblicazione: (2024)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
di: Qiao, Yuxuan, et al.
Pubblicazione: (2024)
di: Qiao, Yuxuan, et al.
Pubblicazione: (2024)
Are Your LLMs Capable of Stable Reasoning?
di: Liu, Junnan, et al.
Pubblicazione: (2024)
di: Liu, Junnan, et al.
Pubblicazione: (2024)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
di: Liu, Shudong, et al.
Pubblicazione: (2025)
di: Liu, Shudong, et al.
Pubblicazione: (2025)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
di: Li, Xin, et al.
Pubblicazione: (2025)
di: Li, Xin, et al.
Pubblicazione: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
di: Dekoninck, Jasper, et al.
Pubblicazione: (2026)
di: Dekoninck, Jasper, et al.
Pubblicazione: (2026)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
di: Lyu, Chengqi, et al.
Pubblicazione: (2025)
di: Lyu, Chengqi, et al.
Pubblicazione: (2025)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
di: Fang, Meng, et al.
Pubblicazione: (2024)
di: Fang, Meng, et al.
Pubblicazione: (2024)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
di: Gu, Yuzhe, et al.
Pubblicazione: (2025)
di: Gu, Yuzhe, et al.
Pubblicazione: (2025)
UrBench: A Comprehensive Benchmark for Evaluating Large Multimodal Models in Multi-View Urban Scenarios
di: Zhou, Baichuan, et al.
Pubblicazione: (2024)
di: Zhou, Baichuan, et al.
Pubblicazione: (2024)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
di: Wang, Yiming, et al.
Pubblicazione: (2025)
di: Wang, Yiming, et al.
Pubblicazione: (2025)
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
di: Li, Lijun, et al.
Pubblicazione: (2024)
di: Li, Lijun, et al.
Pubblicazione: (2024)
PAI-Bench: A Comprehensive Benchmark For Physical AI
di: Zhou, Fengzhe, et al.
Pubblicazione: (2025)
di: Zhou, Fengzhe, et al.
Pubblicazione: (2025)
JMigBench: A Benchmark for Evaluating LLMs on Source Code Migration (Java 8 to Java 11)
di: Amin, Nishil, et al.
Pubblicazione: (2026)
di: Amin, Nishil, et al.
Pubblicazione: (2026)
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
di: Hua, Zhouqi, et al.
Pubblicazione: (2025)
di: Hua, Zhouqi, et al.
Pubblicazione: (2025)
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
di: Ding, Shuangrui, et al.
Pubblicazione: (2026)
di: Ding, Shuangrui, et al.
Pubblicazione: (2026)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
di: Wang, Lei, et al.
Pubblicazione: (2024)
di: Wang, Lei, et al.
Pubblicazione: (2024)
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy
di: Zhan, Shaoxiong, et al.
Pubblicazione: (2025)
di: Zhan, Shaoxiong, et al.
Pubblicazione: (2025)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
di: Chen, Zehui, et al.
Pubblicazione: (2023)
di: Chen, Zehui, et al.
Pubblicazione: (2023)
RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
di: Zhang, Jie, et al.
Pubblicazione: (2025)
di: Zhang, Jie, et al.
Pubblicazione: (2025)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
di: Li, Rongjie, et al.
Pubblicazione: (2024)
di: Li, Rongjie, et al.
Pubblicazione: (2024)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
di: Fang, Xinyu, et al.
Pubblicazione: (2024)
di: Fang, Xinyu, et al.
Pubblicazione: (2024)
Disce aut Deficere: Evaluating LLMs Proficiency on the INVALSI Italian Benchmark
di: Mercorio, Fabio, et al.
Pubblicazione: (2024)
di: Mercorio, Fabio, et al.
Pubblicazione: (2024)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
di: Zou, Chengke, et al.
Pubblicazione: (2024)
di: Zou, Chengke, et al.
Pubblicazione: (2024)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
di: Li, Can, et al.
Pubblicazione: (2025)
di: Li, Can, et al.
Pubblicazione: (2025)
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
di: Yang, Tengchao, et al.
Pubblicazione: (2025)
di: Yang, Tengchao, et al.
Pubblicazione: (2025)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models
di: Li, Wei, et al.
Pubblicazione: (2024)
di: Li, Wei, et al.
Pubblicazione: (2024)
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
di: Glazer, Elliot, et al.
Pubblicazione: (2024)
di: Glazer, Elliot, et al.
Pubblicazione: (2024)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
di: Tong, Yuxuan, et al.
Pubblicazione: (2024)
di: Tong, Yuxuan, et al.
Pubblicazione: (2024)
Rectifying LLM Thought from Lens of Optimization
di: Liu, Junnan, et al.
Pubblicazione: (2025)
di: Liu, Junnan, et al.
Pubblicazione: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Li, Xiaoyuan, et al.
Pubblicazione: (2025)
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
di: Yang, Sihan, et al.
Pubblicazione: (2025)
di: Yang, Sihan, et al.
Pubblicazione: (2025)
Achieving Olympia-Level Geometry Large Language Model Agent via Complexity Boosting Reinforcement Learning
di: Zhao, Haiteng, et al.
Pubblicazione: (2025)
di: Zhao, Haiteng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
di: Wang, Chonghua, et al.
Pubblicazione: (2024) -
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
di: Zhang, Chuyu, et al.
Pubblicazione: (2024) -
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
di: Zhuo, Jingming, et al.
Pubblicazione: (2024) -
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
di: Li, Mo, et al.
Pubblicazione: (2024) -
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
di: Ying, Huaiyuan, et al.
Pubblicazione: (2024)