A Lean Dataset for International Math Olympiad: Small Steps towards Writing Math Proofs for Hard Problems
Fuente:
arXiv
Guardado en:
| Autores principales: | Yousefzadeh, Roozbeh, Cao, Xuenan, Ospanov, Azim |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Advocate for Complete Benchmarks for Formal Reasoning with Formal/Informal Statements and Formal/Informal Proofs
por: Yousefzadeh, Roozbeh, et al.
Publicado: (2025)
por: Yousefzadeh, Roozbeh, et al.
Publicado: (2025)
APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning
por: Ospanov, Azim, et al.
Publicado: (2025)
por: Ospanov, Azim, et al.
Publicado: (2025)
miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
por: Ospanov, Azim, et al.
Publicado: (2025)
por: Ospanov, Azim, et al.
Publicado: (2025)
Towards a Scalable Reference-Free Evaluation of Generative Models
por: Ospanov, Azim, et al.
Publicado: (2024)
por: Ospanov, Azim, et al.
Publicado: (2024)
Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
por: Ospanov, Azim, et al.
Publicado: (2024)
por: Ospanov, Azim, et al.
Publicado: (2024)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
por: Mahdavi, Sadegh, et al.
Publicado: (2025)
por: Mahdavi, Sadegh, et al.
Publicado: (2025)
Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo
por: Feng, Shengyu, et al.
Publicado: (2024)
por: Feng, Shengyu, et al.
Publicado: (2024)
MathWriting: A Dataset For Handwritten Mathematical Expression Recognition
por: Gervais, Philippe, et al.
Publicado: (2024)
por: Gervais, Philippe, et al.
Publicado: (2024)
FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?
por: Ravi, Nikil, et al.
Publicado: (2026)
por: Ravi, Nikil, et al.
Publicado: (2026)
MathChat: Converse to Tackle Challenging Math Problems with LLM Agents
por: Wu, Yiran, et al.
Publicado: (2023)
por: Wu, Yiran, et al.
Publicado: (2023)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
por: Opedal, Andreas, et al.
Publicado: (2024)
por: Opedal, Andreas, et al.
Publicado: (2024)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
por: Pandit, Shrey, et al.
Publicado: (2025)
por: Pandit, Shrey, et al.
Publicado: (2025)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
por: Toshniwal, Shubham, et al.
Publicado: (2024)
por: Toshniwal, Shubham, et al.
Publicado: (2024)
Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error
por: Farnia, Farzan, et al.
Publicado: (2026)
por: Farnia, Farzan, et al.
Publicado: (2026)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
por: Mahabadi, Rabeeh Karimi, et al.
Publicado: (2025)
por: Mahabadi, Rabeeh Karimi, et al.
Publicado: (2025)
Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
por: Albalak, Alon, et al.
Publicado: (2025)
por: Albalak, Alon, et al.
Publicado: (2025)
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
por: Petrov, Ivo, et al.
Publicado: (2025)
por: Petrov, Ivo, et al.
Publicado: (2025)
MegaMath: Pushing the Limits of Open Math Corpora
por: Zhou, Fan, et al.
Publicado: (2025)
por: Zhou, Fan, et al.
Publicado: (2025)
TreeCut: A Synthetic Unanswerable Math Word Problem Dataset for LLM Hallucination Evaluation
por: Ouyang, Jialin
Publicado: (2025)
por: Ouyang, Jialin
Publicado: (2025)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
por: Wang, Zengzhi, et al.
Publicado: (2023)
por: Wang, Zengzhi, et al.
Publicado: (2023)
Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
por: He, Xuan, et al.
Publicado: (2024)
por: He, Xuan, et al.
Publicado: (2024)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
por: Huang, Kaixuan, et al.
Publicado: (2025)
por: Huang, Kaixuan, et al.
Publicado: (2025)
ControlMath: Controllable Data Generation Promotes Math Generalist Models
por: Chen, Nuo, et al.
Publicado: (2024)
por: Chen, Nuo, et al.
Publicado: (2024)
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
por: Li, Chengpeng, et al.
Publicado: (2023)
por: Li, Chengpeng, et al.
Publicado: (2023)
Solving Formal Math Problems by Decomposition and Iterative Reflection
por: Zhou, Yichi, et al.
Publicado: (2025)
por: Zhou, Yichi, et al.
Publicado: (2025)
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
por: Zhang, Renrui, et al.
Publicado: (2024)
por: Zhang, Renrui, et al.
Publicado: (2024)
Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
por: Di, Xinhan, et al.
Publicado: (2025)
por: Di, Xinhan, et al.
Publicado: (2025)
†DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems
por: Nazi, Zabir Al, et al.
Publicado: (2026)
por: Nazi, Zabir Al, et al.
Publicado: (2026)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
por: Liu, Zihan, et al.
Publicado: (2024)
por: Liu, Zihan, et al.
Publicado: (2024)
Augmenting Math Word Problems via Iterative Question Composing
por: Liu, Haoxiong, et al.
Publicado: (2024)
por: Liu, Haoxiong, et al.
Publicado: (2024)
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
por: Li, Junsong, et al.
Publicado: (2025)
por: Li, Junsong, et al.
Publicado: (2025)
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
por: Wang, Peiyi, et al.
Publicado: (2023)
por: Wang, Peiyi, et al.
Publicado: (2023)
A Small Math Model: Recasting Strategy Choice Theory in an LLM-Inspired Architecture
por: Rahman, Roussel, et al.
Publicado: (2025)
por: Rahman, Roussel, et al.
Publicado: (2025)
Conditional Vendi Score: An Information-Theoretic Approach to Diversity Evaluation of Prompt-based Generative Models
por: Jalali, Mohammad, et al.
Publicado: (2024)
por: Jalali, Mohammad, et al.
Publicado: (2024)
ProofSketcher: Hybrid LLM + Lightweight Proof Checker for Reliable Math/Logic Reasoning
por: Kommuru, Kranthi, et al.
Publicado: (2026)
por: Kommuru, Kranthi, et al.
Publicado: (2026)
EasyMath: A 0-shot Math Benchmark for SLMs
por: Karki, Drishya, et al.
Publicado: (2025)
por: Karki, Drishya, et al.
Publicado: (2025)
DOoM: Difficult Olympiads of Math
por: Kuleshov, Ilya, et al.
Publicado: (2025)
por: Kuleshov, Ilya, et al.
Publicado: (2025)
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
por: Singh, Kunal, et al.
Publicado: (2025)
por: Singh, Kunal, et al.
Publicado: (2025)
MathAtlas: A Benchmark for Autoformalization in the Wild
por: Patel, Nilay, et al.
Publicado: (2026)
por: Patel, Nilay, et al.
Publicado: (2026)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
por: Qin, Tian, et al.
Publicado: (2025)
por: Qin, Tian, et al.
Publicado: (2025)
Ejemplares similares
-
Advocate for Complete Benchmarks for Formal Reasoning with Formal/Informal Statements and Formal/Informal Proofs
por: Yousefzadeh, Roozbeh, et al.
Publicado: (2025) -
APOLLO: Automated LLM and Lean Collaboration for Advanced Formal Reasoning
por: Ospanov, Azim, et al.
Publicado: (2025) -
miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
por: Ospanov, Azim, et al.
Publicado: (2025) -
Towards a Scalable Reference-Free Evaluation of Generative Models
por: Ospanov, Azim, et al.
Publicado: (2024) -
Do Vendi Scores Converge with Finite Samples? Truncated Vendi Score for Finite-Sample Convergence Guarantees
por: Ospanov, Azim, et al.
Publicado: (2024)