S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Muzhi, Yang, Chenxu, Si, Qingyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stable Reinforcement Learning for Efficient Reasoning
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
Dynamic Early Exit in Reasoning Models
by: Yang, Chenxu, et al.
Published: (2025)
by: Yang, Chenxu, et al.
Published: (2025)
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
by: Yang, Rubing, et al.
Published: (2025)
by: Yang, Rubing, et al.
Published: (2025)
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
by: Nagle, Alliot, et al.
Published: (2026)
by: Nagle, Alliot, et al.
Published: (2026)
Beyond Greedy Exits: Improved Early Exit Decisions for Risk Control and Reliability
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Federated Learning for Collaborative Inference Systems: The Case of Early Exit Networks
by: Kaplan, Caelin, et al.
Published: (2024)
by: Kaplan, Caelin, et al.
Published: (2024)
E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models
by: Zhang, Shengjun, et al.
Published: (2026)
by: Zhang, Shengjun, et al.
Published: (2026)
ExGRPO: Learning to Reason from Experience
by: Zhan, Runzhe, et al.
Published: (2025)
by: Zhan, Runzhe, et al.
Published: (2025)
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
by: Chen, Benteng, et al.
Published: (2026)
by: Chen, Benteng, et al.
Published: (2026)
Towards Robust Deep Reinforcement Learning against Environmental State Perturbation
by: Wang, Chenxu, et al.
Published: (2025)
by: Wang, Chenxu, et al.
Published: (2025)
Improving Group Fairness in Knowledge Distillation via Laplace Approximation of Early Exits
by: Fasth, Edvin, et al.
Published: (2025)
by: Fasth, Edvin, et al.
Published: (2025)
AEBNAS: Strengthening Exit Branches in Early-Exit Networks through Hardware-Aware Neural Architecture Search
by: Robben, Oscar, et al.
Published: (2025)
by: Robben, Oscar, et al.
Published: (2025)
Early-Exit Neural Networks with Nested Prediction Sets
by: Jazbec, Metod, et al.
Published: (2023)
by: Jazbec, Metod, et al.
Published: (2023)
Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
by: Bereket, Michael, et al.
Published: (2025)
by: Bereket, Michael, et al.
Published: (2025)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
by: Parthasarathi, Prasanna, et al.
Published: (2025)
by: Parthasarathi, Prasanna, et al.
Published: (2025)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Quantifying and Understanding Uncertainty in Large Reasoning Models
by: Li, Yangyi, et al.
Published: (2026)
by: Li, Yangyi, et al.
Published: (2026)
LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning
by: Akgül, Ömer Faruk, et al.
Published: (2025)
by: Akgül, Ömer Faruk, et al.
Published: (2025)
Deep Reinforcement Learning with Task-Adaptive Retrieval via Hypernetwork
by: Jin, Yonggang, et al.
Published: (2023)
by: Jin, Yonggang, et al.
Published: (2023)
Attention Consistency Regularization for Interpretable Early-Exit Neural Networks
by: Zhao, Yanhua
Published: (2026)
by: Zhao, Yanhua
Published: (2026)
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
by: Zheng, Zhi, et al.
Published: (2025)
by: Zheng, Zhi, et al.
Published: (2025)
Hetero-SplitEE: Split Learning of Neural Networks with Early Exits for Heterogeneous IoT Devices
by: Oda, Yuki, et al.
Published: (2025)
by: Oda, Yuki, et al.
Published: (2025)
TableGPT-R1: Advancing Tabular Reasoning Through Reinforcement Learning
by: Yang, Saisai, et al.
Published: (2025)
by: Yang, Saisai, et al.
Published: (2025)
Bridging Efficiency and Safety: Formal Verification of Neural Networks with Early Exits
by: Elboher, Yizhak Yisrael, et al.
Published: (2025)
by: Elboher, Yizhak Yisrael, et al.
Published: (2025)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
by: Yari, Amir Hossein, et al.
Published: (2026)
by: Yari, Amir Hossein, et al.
Published: (2026)
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
by: Monsefi, Amin Karimi, et al.
Published: (2026)
by: Monsefi, Amin Karimi, et al.
Published: (2026)
MC-GRPO: Median-Centered Group Relative Policy Optimization for Small-Rollout Reinforcement Learning
by: Kim, Youngeun
Published: (2026)
by: Kim, Youngeun
Published: (2026)
On-Sensor Convolutional Neural Networks with Early-Exits
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
by: Shalby, Hazem Hesham Yousef, et al.
Published: (2025)
Dynamic Vocabulary Pruning in Early-Exit LLMs
by: Vincenti, Jort, et al.
Published: (2024)
by: Vincenti, Jort, et al.
Published: (2024)
Know What You Don't Know: Selective Prediction for Early Exit DNNs
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
by: Bajpai, Divya Jyoti, et al.
Published: (2025)
Comparative Analysis and Parametric Tuning of PPO, GRPO, and DAPO for LLM Reasoning Enhancement
by: Lian, Yongsheng
Published: (2025)
by: Lian, Yongsheng
Published: (2025)
GRPO is Secretly a Process Reward Model
by: Sullivan, Michael, et al.
Published: (2025)
by: Sullivan, Michael, et al.
Published: (2025)
Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
by: Zhu, Mingkang, et al.
Published: (2025)
by: Zhu, Mingkang, et al.
Published: (2025)
Temporal Decisions: Leveraging Temporal Correlation for Efficient Decisions in Early Exit Neural Networks
by: Sponner, Max, et al.
Published: (2024)
by: Sponner, Max, et al.
Published: (2024)
Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO
by: Yu, Bowen, et al.
Published: (2026)
by: Yu, Bowen, et al.
Published: (2026)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
From Reasoning to Code: GRPO Optimization for Underrepresented Languages
by: Pennino, Federico, et al.
Published: (2025)
by: Pennino, Federico, et al.
Published: (2025)
Mitigating Overthinking in Large Reasoning Models via Difficulty-aware Reinforcement Learning
by: Wan, Qian, et al.
Published: (2026)
by: Wan, Qian, et al.
Published: (2026)
Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models
by: Yang, Chenxu, et al.
Published: (2025)
by: Yang, Chenxu, et al.
Published: (2025)
Multi-Path Collaborative Reasoning via Reinforcement Learning
by: Lv, Jindi, et al.
Published: (2025)
by: Lv, Jindi, et al.
Published: (2025)
Similar Items
-
Stable Reinforcement Learning for Efficient Reasoning
by: Dai, Muzhi, et al.
Published: (2025) -
Dynamic Early Exit in Reasoning Models
by: Yang, Chenxu, et al.
Published: (2025) -
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
by: Yang, Rubing, et al.
Published: (2025) -
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
by: Nagle, Alliot, et al.
Published: (2026) -
Beyond Greedy Exits: Improved Early Exit Decisions for Risk Control and Reliability
by: Bajpai, Divya Jyoti, et al.
Published: (2025)