EternalMath: A Living Benchmark of Frontier Mathematics that Evolves with Human Discovery
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Jicheng, Wang, Guohua, Feng, Xinhua, Liu, Yiming, Hu, Zhichao, Liu, Yuhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
von: Xue, Boyang, et al.
Veröffentlicht: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025)
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
von: Guo, Dadi, et al.
Veröffentlicht: (2026)
von: Guo, Dadi, et al.
Veröffentlicht: (2026)
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
von: Liang, Hao, et al.
Veröffentlicht: (2025)
von: Liang, Hao, et al.
Veröffentlicht: (2025)
UltraLogic: Enhancing LLM Reasoning through Large-Scale Data Synthesis and Bipolar Float Reward
von: Liu, Yile, et al.
Veröffentlicht: (2026)
von: Liu, Yile, et al.
Veröffentlicht: (2026)
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
von: Li, Xin, et al.
Veröffentlicht: (2025)
von: Li, Xin, et al.
Veröffentlicht: (2025)
RoMath: A Mathematical Reasoning Benchmark in Romanian
von: Cosma, Adrian, et al.
Veröffentlicht: (2024)
von: Cosma, Adrian, et al.
Veröffentlicht: (2024)
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
von: Liu, Hongwei, et al.
Veröffentlicht: (2024)
PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation
von: Wei, Hu, et al.
Veröffentlicht: (2025)
von: Wei, Hu, et al.
Veröffentlicht: (2025)
FineMath: A Fine-Grained Mathematical Evaluation Benchmark for Chinese Large Language Models
von: Liu, Yan, et al.
Veröffentlicht: (2024)
von: Liu, Yan, et al.
Veröffentlicht: (2024)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
von: Sun, Yuhong, et al.
Veröffentlicht: (2024)
von: Sun, Yuhong, et al.
Veröffentlicht: (2024)
CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models
von: Liu, Wentao, et al.
Veröffentlicht: (2024)
von: Liu, Wentao, et al.
Veröffentlicht: (2024)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
von: Ma, Jingkun, et al.
Veröffentlicht: (2024)
von: Ma, Jingkun, et al.
Veröffentlicht: (2024)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
von: Guan, Xinyu, et al.
Veröffentlicht: (2025)
von: Guan, Xinyu, et al.
Veröffentlicht: (2025)
MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
von: O'Brien, Dayyán, et al.
Veröffentlicht: (2025)
von: O'Brien, Dayyán, et al.
Veröffentlicht: (2025)
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
von: Zhou, Zihao, et al.
Veröffentlicht: (2024)
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
von: Sobhani, Mahbub E, et al.
Veröffentlicht: (2025)
von: Sobhani, Mahbub E, et al.
Veröffentlicht: (2025)
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
von: Wu, Yanan, et al.
Veröffentlicht: (2024)
von: Wu, Yanan, et al.
Veröffentlicht: (2024)
ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
von: Liu, Hongwei, et al.
Veröffentlicht: (2025)
von: Liu, Hongwei, et al.
Veröffentlicht: (2025)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
MathLearner: A Large Language Model Agent Framework for Learning to Solve Mathematical Problems
von: Xie, Wenbei, et al.
Veröffentlicht: (2024)
von: Xie, Wenbei, et al.
Veröffentlicht: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
von: Fang, Meng, et al.
Veröffentlicht: (2024)
von: Fang, Meng, et al.
Veröffentlicht: (2024)
Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models
von: Shi, Wenhao, et al.
Veröffentlicht: (2024)
von: Shi, Wenhao, et al.
Veröffentlicht: (2024)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
von: Li, Zhuohao, et al.
Veröffentlicht: (2025)
von: Li, Zhuohao, et al.
Veröffentlicht: (2025)
JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
von: Wei, Chengwei, et al.
Veröffentlicht: (2025)
von: Wei, Chengwei, et al.
Veröffentlicht: (2025)
ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2026)
von: Zhao, Raoyuan, et al.
Veröffentlicht: (2026)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving
von: Colle, Vincenzo, et al.
Veröffentlicht: (2025)
von: Colle, Vincenzo, et al.
Veröffentlicht: (2025)
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization
von: Chi, Yizhe, et al.
Veröffentlicht: (2026)
von: Chi, Yizhe, et al.
Veröffentlicht: (2026)
The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark
von: Liu, Hao, et al.
Veröffentlicht: (2026)
von: Liu, Hao, et al.
Veröffentlicht: (2026)
MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical Contexts
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
von: Yang, Shidong, et al.
Veröffentlicht: (2026)
von: Yang, Shidong, et al.
Veröffentlicht: (2026)
SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
von: Manem, Chaitanya, et al.
Veröffentlicht: (2025)
von: Manem, Chaitanya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models
von: Xue, Boyang, et al.
Veröffentlicht: (2025) -
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2025) -
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
von: Guo, Dadi, et al.
Veröffentlicht: (2026) -
MathClean: A Benchmark for Synthetic Mathematical Data Cleaning
von: Liang, Hao, et al.
Veröffentlicht: (2025) -
UltraLogic: Enhancing LLM Reasoning through Large-Scale Data Synthesis and Bipolar Float Reward
von: Liu, Yile, et al.
Veröffentlicht: (2026)