EvolMathEval: Towards Evolvable Benchmarks for Mathematical Reasoning via Evolutionary Testing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Shengbo, Liu, Mingwei, Li, Zike, Li, Anji, Wang, Yanlin, Peng, Xin, Zheng, Zibin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Preliminary Study on the Robustness of Code Generation by Large Language Models
von: Li, Zike, et al.
Veröffentlicht: (2025)
von: Li, Zike, et al.
Veröffentlicht: (2025)
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
von: Duan, Guoliang, et al.
Veröffentlicht: (2025)
von: Duan, Guoliang, et al.
Veröffentlicht: (2025)
KTester: Leveraging Domain and Testing Knowledge for More Effective LLM-based Test Generation
von: Li, Anji, et al.
Veröffentlicht: (2025)
von: Li, Anji, et al.
Veröffentlicht: (2025)
Build-Aware Incremental C-to-Rust Migration via Skeleton-First Translation and Historical Knowledge Reuse
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
von: Dai, Dekun, et al.
Veröffentlicht: (2025)
von: Dai, Dekun, et al.
Veröffentlicht: (2025)
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
von: Luo, Haipeng, et al.
Veröffentlicht: (2023)
von: Luo, Haipeng, et al.
Veröffentlicht: (2023)
AdaDec: A Uncertainty-Guided Lookahead Decoding Framework for LLM-Based Code Generation
von: He, Kaifeng, et al.
Veröffentlicht: (2025)
von: He, Kaifeng, et al.
Veröffentlicht: (2025)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
von: Liang, Linxi, et al.
Veröffentlicht: (2025)
von: Liang, Linxi, et al.
Veröffentlicht: (2025)
CoSQA+: Pioneering the Multi-Choice Code Search Benchmark with Test-Driven Agents
von: Gong, Jing, et al.
Veröffentlicht: (2024)
von: Gong, Jing, et al.
Veröffentlicht: (2024)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
von: Li, Can, et al.
Veröffentlicht: (2025)
von: Li, Can, et al.
Veröffentlicht: (2025)
RustRepoTrans: Repository-level Code Translation Benchmark Targeting Rust
von: Ou, Guangsheng, et al.
Veröffentlicht: (2024)
von: Ou, Guangsheng, et al.
Veröffentlicht: (2024)
Are Decoder-Only Large Language Models the Silver Bullet for Code Search?
von: Chen, Yuxuan, et al.
Veröffentlicht: (2024)
von: Chen, Yuxuan, et al.
Veröffentlicht: (2024)
Generating High-Quality Datasets for Code Editing via Open-Source Language Models
von: Zhang, Zekai, et al.
Veröffentlicht: (2025)
von: Zhang, Zekai, et al.
Veröffentlicht: (2025)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
von: Li, Chengpeng, et al.
Veröffentlicht: (2024)
von: Li, Chengpeng, et al.
Veröffentlicht: (2024)
RoMath: A Mathematical Reasoning Benchmark in Romanian
von: Cosma, Adrian, et al.
Veröffentlicht: (2024)
von: Cosma, Adrian, et al.
Veröffentlicht: (2024)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
von: Ma, Jingkun, et al.
Veröffentlicht: (2024)
von: Ma, Jingkun, et al.
Veröffentlicht: (2024)
Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code
von: He, Kaifeng, et al.
Veröffentlicht: (2026)
von: He, Kaifeng, et al.
Veröffentlicht: (2026)
A Historical Trajectory Assisted Optimization Method for Zeroth-Order Federated Learning
von: Wu, Chenlin, et al.
Veröffentlicht: (2024)
von: Wu, Chenlin, et al.
Veröffentlicht: (2024)
GraCoRe: Benchmarking Graph Comprehension and Complex Reasoning in Large Language Models
von: Yuan, Zike, et al.
Veröffentlicht: (2024)
von: Yuan, Zike, et al.
Veröffentlicht: (2024)
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
von: Glazer, Elliot, et al.
Veröffentlicht: (2024)
von: Glazer, Elliot, et al.
Veröffentlicht: (2024)
Direct Preference-Based Evolutionary Multi-Objective Optimization with Dueling Bandit
von: Huang, Tian, et al.
Veröffentlicht: (2023)
von: Huang, Tian, et al.
Veröffentlicht: (2023)
Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models
von: Wang, Yanlin, et al.
Veröffentlicht: (2024)
von: Wang, Yanlin, et al.
Veröffentlicht: (2024)
MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
von: Liang, Zhenwen, et al.
Veröffentlicht: (2024)
Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents
von: Li, Junkai, et al.
Veröffentlicht: (2024)
von: Li, Junkai, et al.
Veröffentlicht: (2024)
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
von: Lei, Yingtie, et al.
Veröffentlicht: (2026)
von: Lei, Yingtie, et al.
Veröffentlicht: (2026)
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
von: Wu, Yanan, et al.
Veröffentlicht: (2024)
von: Wu, Yanan, et al.
Veröffentlicht: (2024)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
EvolVE: Evolutionary Search for LLM-based Verilog Generation and Optimization
von: Hsin, Wei-Po, et al.
Veröffentlicht: (2026)
von: Hsin, Wei-Po, et al.
Veröffentlicht: (2026)
Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning
von: Zhang, Yanzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Yanzhi, et al.
Veröffentlicht: (2025)
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
von: Alshammari, Shaden, et al.
Veröffentlicht: (2026)
von: Alshammari, Shaden, et al.
Veröffentlicht: (2026)
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
von: Qiao, Runqi, et al.
Veröffentlicht: (2025)
von: Qiao, Runqi, et al.
Veröffentlicht: (2025)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
von: Zheng, Dewu, et al.
Veröffentlicht: (2024)
von: Zheng, Dewu, et al.
Veröffentlicht: (2024)
Knowledge-Graph-Driven Data Synthesis for Low-Resource Software Development: A HarmonyOS Case Study
von: Liu, Mingwei, et al.
Veröffentlicht: (2025)
von: Liu, Mingwei, et al.
Veröffentlicht: (2025)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
von: Yu, Zhouliang, et al.
Veröffentlicht: (2025)
von: Yu, Zhouliang, et al.
Veröffentlicht: (2025)
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
von: Chen, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyang, et al.
Veröffentlicht: (2025)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning
von: Li, Xuchen, et al.
Veröffentlicht: (2026)
von: Li, Xuchen, et al.
Veröffentlicht: (2026)
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
von: Shao, Zhihong, et al.
Veröffentlicht: (2025)
von: Shao, Zhihong, et al.
Veröffentlicht: (2025)
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task
von: Yan, Yuchen, et al.
Veröffentlicht: (2025)
von: Yan, Yuchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Preliminary Study on the Robustness of Code Generation by Large Language Models
von: Li, Zike, et al.
Veröffentlicht: (2025) -
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
von: Duan, Guoliang, et al.
Veröffentlicht: (2025) -
KTester: Leveraging Domain and Testing Knowledge for More Effective LLM-based Test Generation
von: Li, Anji, et al.
Veröffentlicht: (2025) -
Build-Aware Incremental C-to-Rust Migration via Skeleton-First Translation and Historical Knowledge Reuse
von: Wang, Shengbo, et al.
Veröffentlicht: (2026) -
FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
von: Dai, Dekun, et al.
Veröffentlicht: (2025)