A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yoshihara, Hiroshi, Yamaguchi, Taiki, Inoue, Yuichi |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
par: Peng, Ruiying, et autres
Publié: (2026)
par: Peng, Ruiying, et autres
Publié: (2026)
A Recipe for Stable Offline Multi-agent Reinforcement Learning
par: Lee, Dongsu, et autres
Publié: (2026)
par: Lee, Dongsu, et autres
Publié: (2026)
Reset & Distill: A Recipe for Overcoming Negative Transfer in Continual Reinforcement Learning
par: Ahn, Hongjoon, et autres
Publié: (2024)
par: Ahn, Hongjoon, et autres
Publié: (2024)
PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning
par: Wang, Ruheng, et autres
Publié: (2025)
par: Wang, Ruheng, et autres
Publié: (2025)
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
par: Koh, Woosung, et autres
Publié: (2026)
par: Koh, Woosung, et autres
Publié: (2026)
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
par: Gan, Xingwei, et autres
Publié: (2026)
par: Gan, Xingwei, et autres
Publié: (2026)
Debunk the Myth of SFT Generalization
par: Lin, Xiaofeng, et autres
Publié: (2025)
par: Lin, Xiaofeng, et autres
Publié: (2025)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
par: Kim, Junhan, et autres
Publié: (2026)
par: Kim, Junhan, et autres
Publié: (2026)
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
par: Zheng, Chujie, et autres
Publié: (2025)
par: Zheng, Chujie, et autres
Publié: (2025)
Unsupervised Learning in Echo State Networks for Input Reconstruction
par: Yamada, Taiki, et autres
Publié: (2025)
par: Yamada, Taiki, et autres
Publié: (2025)
TSSR: Two-Stage Swap-Reward-Driven Reinforcement Learning for Character-Level SMILES Generation
par: Levine, Jacob Ede, et autres
Publié: (2026)
par: Levine, Jacob Ede, et autres
Publié: (2026)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
par: Feng, Steven, et autres
Publié: (2024)
par: Feng, Steven, et autres
Publié: (2024)
Machine Learning-Based Quantification of Vesicoureteral Reflux with Enhancing Accuracy and Efficiency
par: Alqaraleh, Muhyeeddin, et autres
Publié: (2025)
par: Alqaraleh, Muhyeeddin, et autres
Publié: (2025)
Pentest-R1: Towards Autonomous Penetration Testing Reasoning Optimized via Two-Stage Reinforcement Learning
par: Kong, He, et autres
Publié: (2025)
par: Kong, He, et autres
Publié: (2025)
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
par: Gong, Zixuan, et autres
Publié: (2025)
par: Gong, Zixuan, et autres
Publié: (2025)
DISPO: Enhancing Training Efficiency and Stability in Reinforcement Learning for Large Language Model Mathematical Reasoning
par: Karaman, Batuhan K., et autres
Publié: (2026)
par: Karaman, Batuhan K., et autres
Publié: (2026)
RL Fine-Tuning Heals OOD Forgetting in SFT
par: Jin, Hangzhan, et autres
Publié: (2025)
par: Jin, Hangzhan, et autres
Publié: (2025)
Two-Stage Regularization-Based Structured Pruning for LLMs
par: Feng, Mingkuan, et autres
Publié: (2025)
par: Feng, Mingkuan, et autres
Publié: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
par: Sun, Yiyou, et autres
Publié: (2025)
par: Sun, Yiyou, et autres
Publié: (2025)
Two-Stage Learned Decomposition for Scalable Routing on Multigraphs
par: Rydin, Filip, et autres
Publié: (2026)
par: Rydin, Filip, et autres
Publié: (2026)
A Practical Introduction to Deep Reinforcement Learning
par: Sun, Yinghan, et autres
Publié: (2025)
par: Sun, Yinghan, et autres
Publié: (2025)
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
par: Sandri, Fabrizio, et autres
Publié: (2025)
par: Sandri, Fabrizio, et autres
Publié: (2025)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
par: He, Shenghua, et autres
Publié: (2025)
par: He, Shenghua, et autres
Publié: (2025)
The Two-Stage Decision-Sampling Hypothesis: Understanding the Emergence of Self-Reflection in RL-Trained LLMs
par: Zhao, Zibo, et autres
Publié: (2026)
par: Zhao, Zibo, et autres
Publié: (2026)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
par: Kang, Feiyang, et autres
Publié: (2025)
par: Kang, Feiyang, et autres
Publié: (2025)
The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
par: Zimmer, Max, et autres
Publié: (2026)
par: Zimmer, Max, et autres
Publié: (2026)
Recipes for Pre-training LLMs with MXFP8
par: Mishra, Asit, et autres
Publié: (2025)
par: Mishra, Asit, et autres
Publié: (2025)
Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
par: Zhang, Lunjun, et autres
Publié: (2026)
par: Zhang, Lunjun, et autres
Publié: (2026)
Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training Stages
par: Ma, Guozheng, et autres
Publié: (2023)
par: Ma, Guozheng, et autres
Publié: (2023)
Brain-Inspired Two-Stage Approach: Enhancing Mathematical Reasoning by Imitating Human Thought Processes
par: Chen, Yezeng, et autres
Publié: (2024)
par: Chen, Yezeng, et autres
Publié: (2024)
Learning Scenario Reduction for Two-Stage Robust Optimization with Discrete Uncertainty
par: Lin, Tianjue, et autres
Publié: (2026)
par: Lin, Tianjue, et autres
Publié: (2026)
Sensor Calibration Model Balancing Accuracy, Real-time, and Efficiency
par: Yun, Jinyong, et autres
Publié: (2025)
par: Yun, Jinyong, et autres
Publié: (2025)
Step Rejection Fine-Tuning: A Practical Distillation Recipe
par: Slinko, Igor, et autres
Publié: (2026)
par: Slinko, Igor, et autres
Publié: (2026)
The Art of Scaling Reinforcement Learning Compute for LLMs
par: Khatri, Devvrit, et autres
Publié: (2025)
par: Khatri, Devvrit, et autres
Publié: (2025)
RLAE: Reinforcement Learning-Assisted Ensemble for LLMs
par: Fu, Yuqian, et autres
Publié: (2025)
par: Fu, Yuqian, et autres
Publié: (2025)
Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual Learning
par: Qiu, Haomiao, et autres
Publié: (2025)
par: Qiu, Haomiao, et autres
Publié: (2025)
The Recipe Matters More Than the Kitchen:Mathematical Foundations of the AI Weather Prediction Pipeline
par: Garg, Piyush, et autres
Publié: (2026)
par: Garg, Piyush, et autres
Publié: (2026)
FedDRL: A Trustworthy Federated Learning Model Fusion Method Based on Staged Reinforcement Learning
par: Chen, Leiming, et autres
Publié: (2023)
par: Chen, Leiming, et autres
Publié: (2023)
Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning
par: Bozkurt, Alper Kamil, et autres
Publié: (2026)
par: Bozkurt, Alper Kamil, et autres
Publié: (2026)
SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models
par: Kim, Gyuhak, et autres
Publié: (2025)
par: Kim, Gyuhak, et autres
Publié: (2025)
Documents similaires
-
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
par: Peng, Ruiying, et autres
Publié: (2026) -
A Recipe for Stable Offline Multi-agent Reinforcement Learning
par: Lee, Dongsu, et autres
Publié: (2026) -
Reset & Distill: A Recipe for Overcoming Negative Transfer in Continual Reinforcement Learning
par: Ahn, Hongjoon, et autres
Publié: (2024) -
PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning
par: Wang, Ruheng, et autres
Publié: (2025) -
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
par: Koh, Woosung, et autres
Publié: (2026)