MegaMath: Pushing the Limits of Open Math Corpora
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Fan, Wang, Zengzhi, Ranjan, Nikhil, Cheng, Zhoujun, Tang, Liping, He, Guowei, Liu, Zhengzhong, Xing, Eric P. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
von: Fan, Run-Ze, et al.
Veröffentlicht: (2025)
von: Fan, Run-Ze, et al.
Veröffentlicht: (2025)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
von: Wang, Zengzhi, et al.
Veröffentlicht: (2023)
von: Wang, Zengzhi, et al.
Veröffentlicht: (2023)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
von: Shao, Zhihong, et al.
Veröffentlicht: (2024)
von: Shao, Zhihong, et al.
Veröffentlicht: (2024)
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2024)
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2024)
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
von: Liu, Xianyang, et al.
Veröffentlicht: (2025)
von: Liu, Xianyang, et al.
Veröffentlicht: (2025)
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2025)
von: Tian, Shi-Yu, et al.
Veröffentlicht: (2025)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2024)
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
von: Fang, Meng, et al.
Veröffentlicht: (2024)
von: Fang, Meng, et al.
Veröffentlicht: (2024)
Orca-Math: Unlocking the potential of SLMs in Grade School Math
von: Mitra, Arindam, et al.
Veröffentlicht: (2024)
von: Mitra, Arindam, et al.
Veröffentlicht: (2024)
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
von: Li, Chengpeng, et al.
Veröffentlicht: (2023)
von: Li, Chengpeng, et al.
Veröffentlicht: (2023)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
MathBuddy: A Multimodal System for Affective Math Tutoring
von: Kar, Debanjana, et al.
Veröffentlicht: (2025)
von: Kar, Debanjana, et al.
Veröffentlicht: (2025)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
von: Chen, Nuo, et al.
Veröffentlicht: (2024)
von: Chen, Nuo, et al.
Veröffentlicht: (2024)
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese
von: Xu, Liang, et al.
Veröffentlicht: (2024)
von: Xu, Liang, et al.
Veröffentlicht: (2024)
Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward Passes
von: Christ, Bryan R., et al.
Veröffentlicht: (2024)
von: Christ, Bryan R., et al.
Veröffentlicht: (2024)
Let's Verify Math Questions Step by Step
von: Shen, Chengyu, et al.
Veröffentlicht: (2025)
von: Shen, Chengyu, et al.
Veröffentlicht: (2025)
AlphaMath Almost Zero: Process Supervision without Process
von: Chen, Guoxin, et al.
Veröffentlicht: (2024)
von: Chen, Guoxin, et al.
Veröffentlicht: (2024)
ChatGPT as a Math Questioner? Evaluating ChatGPT on Generating Pre-university Math Questions
von: Van Long, Phuoc Pham, et al.
Veröffentlicht: (2023)
von: Van Long, Phuoc Pham, et al.
Veröffentlicht: (2023)
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange
von: Satpute, Ankit, et al.
Veröffentlicht: (2024)
von: Satpute, Ankit, et al.
Veröffentlicht: (2024)
REAMS: Reasoning Enhanced Algorithm for Maths Solving
von: Singh, Eishkaran, et al.
Veröffentlicht: (2025)
von: Singh, Eishkaran, et al.
Veröffentlicht: (2025)
InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
von: Ying, Huaiyuan, et al.
Veröffentlicht: (2024)
Adversarial Math Word Problem Generation
von: Xie, Roy, et al.
Veröffentlicht: (2024)
von: Xie, Roy, et al.
Veröffentlicht: (2024)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
Case-Based or Rule-Based: How Do Transformers Do the Math?
von: Hu, Yi, et al.
Veröffentlicht: (2024)
von: Hu, Yi, et al.
Veröffentlicht: (2024)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
von: Qin, Tian, et al.
Veröffentlicht: (2025)
von: Qin, Tian, et al.
Veröffentlicht: (2025)
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
von: Lu, Zimu, et al.
Veröffentlicht: (2024)
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
von: Pang, Bo, et al.
Veröffentlicht: (2025)
von: Pang, Bo, et al.
Veröffentlicht: (2025)
PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models
von: Peng, Shuai, et al.
Veröffentlicht: (2024)
von: Peng, Shuai, et al.
Veröffentlicht: (2024)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2025)
von: Mahabadi, Rabeeh Karimi, et al.
Veröffentlicht: (2025)
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
von: Albalak, Alon, et al.
Veröffentlicht: (2025)
von: Albalak, Alon, et al.
Veröffentlicht: (2025)
Self-Consistency Boosts Calibration for Math Reasoning
von: Wang, Ante, et al.
Veröffentlicht: (2024)
von: Wang, Ante, et al.
Veröffentlicht: (2024)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
von: Tong, Yuxuan, et al.
Veröffentlicht: (2024)
von: Tong, Yuxuan, et al.
Veröffentlicht: (2024)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
von: Li, Hang, et al.
Veröffentlicht: (2024)
von: Li, Hang, et al.
Veröffentlicht: (2024)
EasyMath: A 0-shot Math Benchmark for SLMs
von: Karki, Drishya, et al.
Veröffentlicht: (2025)
von: Karki, Drishya, et al.
Veröffentlicht: (2025)
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
von: Wang, Zengzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zengzhi, et al.
Veröffentlicht: (2025)
CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models
von: Li, Zhong-Zhi, et al.
Veröffentlicht: (2024)
von: Li, Zhong-Zhi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
von: Fan, Run-Ze, et al.
Veröffentlicht: (2025) -
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
von: Wang, Zengzhi, et al.
Veröffentlicht: (2023) -
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
von: Shao, Zhihong, et al.
Veröffentlicht: (2024) -
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
von: Toshniwal, Shubham, et al.
Veröffentlicht: (2024) -
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
von: Liu, Xianyang, et al.
Veröffentlicht: (2025)