Stable Reinforcement Learning for Efficient Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Muzhi, Liu, Shixuan, Si, Qingyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
Latent-Space Contrastive Reinforcement Learning for Stable and Efficient LLM Reasoning
by: Shan, Lianlei, et al.
Published: (2026)
by: Shan, Lianlei, et al.
Published: (2026)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
by: Wu, Wenbo, et al.
Published: (2025)
by: Wu, Wenbo, et al.
Published: (2025)
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
by: Liu, Ziru, et al.
Published: (2025)
by: Liu, Ziru, et al.
Published: (2025)
TableGPT-R1: Advancing Tabular Reasoning Through Reinforcement Learning
by: Yang, Saisai, et al.
Published: (2025)
by: Yang, Saisai, et al.
Published: (2025)
Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
by: Melo, Luckeciano C., et al.
Published: (2025)
by: Melo, Luckeciano C., et al.
Published: (2025)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
by: Zhang, Hongming, et al.
Published: (2023)
by: Zhang, Hongming, et al.
Published: (2023)
Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
by: Xiang, Violet, et al.
Published: (2025)
by: Xiang, Violet, et al.
Published: (2025)
Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games
by: Ota, Kazuki, et al.
Published: (2026)
by: Ota, Kazuki, et al.
Published: (2026)
Guiding Diffusion Models with Reinforcement Learning for Stable Molecule Generation
by: Zhou, Zhijian, et al.
Published: (2025)
by: Zhou, Zhijian, et al.
Published: (2025)
EXPO: Stable Reinforcement Learning with Expressive Policies
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Conformal Symplectic Optimization for Stable Reinforcement Learning
by: Lyu, Yao, et al.
Published: (2024)
by: Lyu, Yao, et al.
Published: (2024)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
by: Zheng, Haizhong, et al.
Published: (2025)
by: Zheng, Haizhong, et al.
Published: (2025)
CoPRIS: Efficient and Stable Reinforcement Learning via Concurrency-Controlled Partial Rollout with Importance Sampling
by: Qu, Zekai, et al.
Published: (2025)
by: Qu, Zekai, et al.
Published: (2025)
SATURN: SAT-based Reinforcement Learning to Unleash LLMs Reasoning
by: Liu, Huanyu, et al.
Published: (2025)
by: Liu, Huanyu, et al.
Published: (2025)
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
by: Yue, Bo, et al.
Published: (2024)
by: Yue, Bo, et al.
Published: (2024)
CPGD: Toward Stable Rule-based Reinforcement Learning for Language Models
by: Liu, Zongkai, et al.
Published: (2025)
by: Liu, Zongkai, et al.
Published: (2025)
Adaptive Learning of the Latent Space of Wasserstein Generative Adversarial Networks
by: Qiu, Yixuan, et al.
Published: (2024)
by: Qiu, Yixuan, et al.
Published: (2024)
Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security
by: Dai, Muzhi, et al.
Published: (2025)
by: Dai, Muzhi, et al.
Published: (2025)
ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning
by: Sun, Zhishen, et al.
Published: (2026)
by: Sun, Zhishen, et al.
Published: (2026)
In-context Exploration-Exploitation for Reinforcement Learning
by: Dai, Zhenwen, et al.
Published: (2024)
by: Dai, Zhenwen, et al.
Published: (2024)
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
by: Wang, Shenzhi, et al.
Published: (2025)
by: Wang, Shenzhi, et al.
Published: (2025)
ARCLE: The Abstraction and Reasoning Corpus Learning Environment for Reinforcement Learning
by: Lee, Hosung, et al.
Published: (2024)
by: Lee, Hosung, et al.
Published: (2024)
When is Offline Policy Selection Sample Efficient for Reinforcement Learning?
by: Liu, Vincent, et al.
Published: (2023)
by: Liu, Vincent, et al.
Published: (2023)
GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
Multi-Path Collaborative Reasoning via Reinforcement Learning
by: Lv, Jindi, et al.
Published: (2025)
by: Lv, Jindi, et al.
Published: (2025)
Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
by: Pasand, Ali Saheb, et al.
Published: (2026)
by: Pasand, Ali Saheb, et al.
Published: (2026)
Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry
by: Zhang, Guoxi, et al.
Published: (2026)
by: Zhang, Guoxi, et al.
Published: (2026)
HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing
by: Du, Chengyu, et al.
Published: (2026)
by: Du, Chengyu, et al.
Published: (2026)
Minimax Optimal and Computationally Efficient Algorithms for Distributionally Robust Offline Reinforcement Learning
by: Liu, Zhishuai, et al.
Published: (2024)
by: Liu, Zhishuai, et al.
Published: (2024)
Are Large Language Models Table-based Fact-Checkers?
by: Zhang, Hanwen, et al.
Published: (2024)
by: Zhang, Hanwen, et al.
Published: (2024)
Efficient Rectification of Neuro-Symbolic Reasoning Inconsistencies by Abductive Reflection
by: Hu, Wen-Chao, et al.
Published: (2024)
by: Hu, Wen-Chao, et al.
Published: (2024)
Spectral Representation-based Reinforcement Learning
by: Gao, Chenxiao, et al.
Published: (2025)
by: Gao, Chenxiao, et al.
Published: (2025)
SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement Learning
by: Huang, Tairan, et al.
Published: (2025)
by: Huang, Tairan, et al.
Published: (2025)
Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
by: Sun, Zexu, et al.
Published: (2025)
by: Sun, Zexu, et al.
Published: (2025)
G1: Teaching LLMs to Reason on Graphs with Reinforcement Learning
by: Guo, Xiaojun, et al.
Published: (2025)
by: Guo, Xiaojun, et al.
Published: (2025)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
by: Liu, Qihao, et al.
Published: (2025)
by: Liu, Qihao, et al.
Published: (2025)
Causal Information Prioritization for Efficient Reinforcement Learning
by: Cao, Hongye, et al.
Published: (2025)
by: Cao, Hongye, et al.
Published: (2025)
Efficient Reinforcement Learning in Probabilistic Reward Machines
by: Lin, Xiaofeng, et al.
Published: (2024)
by: Lin, Xiaofeng, et al.
Published: (2024)
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
by: Feng, Lang, et al.
Published: (2026)
by: Feng, Lang, et al.
Published: (2026)
Similar Items
-
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
by: Dai, Muzhi, et al.
Published: (2025) -
Latent-Space Contrastive Reinforcement Learning for Stable and Efficient LLM Reasoning
by: Shan, Lianlei, et al.
Published: (2026) -
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
by: Wu, Wenbo, et al.
Published: (2025) -
GHPO: Adaptive Guidance for Stable and Efficient LLM Reinforcement Learning
by: Liu, Ziru, et al.
Published: (2025) -
TableGPT-R1: Advancing Tabular Reasoning Through Reinforcement Learning
by: Yang, Saisai, et al.
Published: (2025)