Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Xuan, Yin, Da, Peng, Nanyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
di: Chen, Huayu, et al.
Pubblicazione: (2025)
di: Chen, Huayu, et al.
Pubblicazione: (2025)
Benchmarking Large Language Models for Math Reasoning Tasks
di: Seßler, Kathrin, et al.
Pubblicazione: (2024)
di: Seßler, Kathrin, et al.
Pubblicazione: (2024)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
di: Huang, Kaixuan, et al.
Pubblicazione: (2025)
FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline
di: Seegmiller, Parker, et al.
Pubblicazione: (2025)
di: Seegmiller, Parker, et al.
Pubblicazione: (2025)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
di: Liu, Wei, et al.
Pubblicazione: (2023)
di: Liu, Wei, et al.
Pubblicazione: (2023)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
di: Pandit, Shrey, et al.
Pubblicazione: (2025)
MathChat: Converse to Tackle Challenging Math Problems with LLM Agents
di: Wu, Yiran, et al.
Pubblicazione: (2023)
di: Wu, Yiran, et al.
Pubblicazione: (2023)
Detecting Machine-Generated Long-Form Content with Latent-Space Variables
di: Tian, Yufei, et al.
Pubblicazione: (2024)
di: Tian, Yufei, et al.
Pubblicazione: (2024)
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
di: Zhou, Ruiyang, et al.
Pubblicazione: (2025)
di: Zhou, Ruiyang, et al.
Pubblicazione: (2025)
How to Make Large Language Models Generate 100% Valid Molecules?
di: Tao, Wen, et al.
Pubblicazione: (2025)
di: Tao, Wen, et al.
Pubblicazione: (2025)
LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring
di: Lee, Unggi, et al.
Pubblicazione: (2026)
di: Lee, Unggi, et al.
Pubblicazione: (2026)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
di: Razin, Noam, et al.
Pubblicazione: (2025)
di: Razin, Noam, et al.
Pubblicazione: (2025)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
di: Liu, Haolin, et al.
Pubblicazione: (2026)
di: Liu, Haolin, et al.
Pubblicazione: (2026)
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
di: Li, Chengpeng, et al.
Pubblicazione: (2023)
di: Li, Chengpeng, et al.
Pubblicazione: (2023)
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
di: Chen, Yang, et al.
Pubblicazione: (2025)
di: Chen, Yang, et al.
Pubblicazione: (2025)
Inference-Time Rethinking with Latent Thought Vectors for Math Reasoning
di: Kong, Deqian, et al.
Pubblicazione: (2026)
di: Kong, Deqian, et al.
Pubblicazione: (2026)
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
di: Liu, Zihan, et al.
Pubblicazione: (2024)
di: Liu, Zihan, et al.
Pubblicazione: (2024)
†DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems
di: Nazi, Zabir Al, et al.
Pubblicazione: (2026)
di: Nazi, Zabir Al, et al.
Pubblicazione: (2026)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
di: Roy, Tiasa Singha, et al.
Pubblicazione: (2025)
DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning
di: Parekh, Tanmay, et al.
Pubblicazione: (2025)
di: Parekh, Tanmay, et al.
Pubblicazione: (2025)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
di: Zhou, Yu, et al.
Pubblicazione: (2025)
di: Zhou, Yu, et al.
Pubblicazione: (2025)
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
di: Parashar, Shubham, et al.
Pubblicazione: (2025)
di: Parashar, Shubham, et al.
Pubblicazione: (2025)
On the Loss of Context-awareness in General Instruction Fine-tuning
di: Wang, Yihan, et al.
Pubblicazione: (2024)
di: Wang, Yihan, et al.
Pubblicazione: (2024)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
di: Yu, Yongcan, et al.
Pubblicazione: (2026)
di: Yu, Yongcan, et al.
Pubblicazione: (2026)
DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning
di: Li, Chengpeng, et al.
Pubblicazione: (2024)
di: Li, Chengpeng, et al.
Pubblicazione: (2024)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
di: Yang, Cheng-Fu, et al.
Pubblicazione: (2024)
di: Yang, Cheng-Fu, et al.
Pubblicazione: (2024)
The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs
di: Bandarkar, Lucas, et al.
Pubblicazione: (2025)
di: Bandarkar, Lucas, et al.
Pubblicazione: (2025)
Language Models Do Hard Arithmetic Tasks Easily and Hardly Do Easy Arithmetic Tasks
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
di: Gambardella, Andrew, et al.
Pubblicazione: (2024)
Laying the Foundation First? Investigating the Generalization from Atomic Skills to Complex Reasoning Tasks
di: Huang, Yuncheng, et al.
Pubblicazione: (2024)
di: Huang, Yuncheng, et al.
Pubblicazione: (2024)
MegaMath: Pushing the Limits of Open Math Corpora
di: Zhou, Fan, et al.
Pubblicazione: (2025)
di: Zhou, Fan, et al.
Pubblicazione: (2025)
RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
di: Setlur, Amrith, et al.
Pubblicazione: (2024)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
di: Tang, Zhengyang, et al.
Pubblicazione: (2024)
di: Tang, Zhengyang, et al.
Pubblicazione: (2024)
REAL Sampling: Boosting Factuality and Diversity of Open-Ended Generation via Asymptotic Entropy
di: Chang, Haw-Shiuan, et al.
Pubblicazione: (2024)
di: Chang, Haw-Shiuan, et al.
Pubblicazione: (2024)
Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning
di: Paul, Bidyarthi, et al.
Pubblicazione: (2025)
di: Paul, Bidyarthi, et al.
Pubblicazione: (2025)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
di: Qin, Tian, et al.
Pubblicazione: (2025)
di: Qin, Tian, et al.
Pubblicazione: (2025)
FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback
di: Wu, Xueqing, et al.
Pubblicazione: (2025)
di: Wu, Xueqing, et al.
Pubblicazione: (2025)
Intrinsic Language-Guided Exploration for Complex Long-Horizon Robotic Manipulation Tasks
di: Triantafyllidis, Eleftherios, et al.
Pubblicazione: (2023)
di: Triantafyllidis, Eleftherios, et al.
Pubblicazione: (2023)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
di: Chen, Zui, et al.
Pubblicazione: (2024)
di: Chen, Zui, et al.
Pubblicazione: (2024)
PCL-Reasoner-V1.5: Advancing Math Reasoning with Offline Reinforcement Learning
di: Lu, Yao, et al.
Pubblicazione: (2026)
di: Lu, Yao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
di: Chen, Huayu, et al.
Pubblicazione: (2025) -
Benchmarking Large Language Models for Math Reasoning Tasks
di: Seßler, Kathrin, et al.
Pubblicazione: (2024) -
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
di: Huang, Kaixuan, et al.
Pubblicazione: (2025) -
FLAMES: Improving LLM Math Reasoning via a Fine-Grained Analysis of the Data Synthesis Pipeline
di: Seegmiller, Parker, et al.
Pubblicazione: (2025) -
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
di: Liu, Wei, et al.
Pubblicazione: (2023)