Gespeichert in:
| 1. Verfasser: | Kim, Youngeun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.22582 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GRAPH-GRPO-LEX: Contract Graph Modeling and Reinforcement Learning with Group Relative Policy Optimization
von: Dechtiar, Moriya, et al.
Veröffentlicht: (2025)
von: Dechtiar, Moriya, et al.
Veröffentlicht: (2025)
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering
von: Cai, Yuzhu, et al.
Veröffentlicht: (2026)
von: Cai, Yuzhu, et al.
Veröffentlicht: (2026)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
von: Wang, Jingyi, et al.
Veröffentlicht: (2026)
EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
Train Less, Learn More: Adaptive Efficient Rollout Optimization for Group-Based Reinforcement Learning
von: Zhang, Zhi, et al.
Veröffentlicht: (2026)
von: Zhang, Zhi, et al.
Veröffentlicht: (2026)
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
von: Zhou, Renping, et al.
Veröffentlicht: (2025)
IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
von: Mundada, Gagan, et al.
Veröffentlicht: (2026)
von: Mundada, Gagan, et al.
Veröffentlicht: (2026)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2025)
SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization
von: Zheng, Zhi, et al.
Veröffentlicht: (2025)
von: Zheng, Zhi, et al.
Veröffentlicht: (2025)
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
von: Fang, Yangyi, et al.
Veröffentlicht: (2026)
von: Fang, Yangyi, et al.
Veröffentlicht: (2026)
EchoRL: Reinforcement Learning via Rollout Echoing
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
von: Lu, Xiaodong, et al.
Veröffentlicht: (2026)
GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
von: Zhang, Han, et al.
Veröffentlicht: (2025)
von: Zhang, Han, et al.
Veröffentlicht: (2025)
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
von: Yao, Chaorui, et al.
Veröffentlicht: (2025)
von: Yao, Chaorui, et al.
Veröffentlicht: (2025)
NGRPO: Negative-enhanced Group Relative Policy Optimization
von: Nan, Gongrui, et al.
Veröffentlicht: (2025)
von: Nan, Gongrui, et al.
Veröffentlicht: (2025)
Guided Cooperation in Hierarchical Reinforcement Learning via Model-based Rollout
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
von: Wang, Haoran, et al.
Veröffentlicht: (2023)
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
von: Sane, Soham
Veröffentlicht: (2025)
von: Sane, Soham
Veröffentlicht: (2025)
APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
von: Zhou, Yuzhen, et al.
Veröffentlicht: (2025)
von: Zhou, Yuzhen, et al.
Veröffentlicht: (2025)
Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
F-GRPO: Don't Let Your Policy Learn the Obvious and Forget the Rare
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
von: Plyusov, Daniil, et al.
Veröffentlicht: (2026)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning
von: Wang, Yujie, et al.
Veröffentlicht: (2026)
von: Wang, Yujie, et al.
Veröffentlicht: (2026)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
von: Heuillet, Maxime, et al.
Veröffentlicht: (2025)
A Unified Framework for Rethinking Policy Divergence Measures in GRPO
von: Wu, Qingyuan, et al.
Veröffentlicht: (2026)
von: Wu, Qingyuan, et al.
Veröffentlicht: (2026)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
Superior Computer Chess with Model Predictive Control, Reinforcement Learning, and Rollout
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
von: Gundawar, Atharva, et al.
Veröffentlicht: (2024)
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning
von: Xu, Yuhang, et al.
Veröffentlicht: (2026)
von: Xu, Yuhang, et al.
Veröffentlicht: (2026)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2025)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
von: Li, Gengsheng, et al.
Veröffentlicht: (2026)
von: Li, Gengsheng, et al.
Veröffentlicht: (2026)
Information-Consistent Language Model Recommendations through Group Relative Policy Optimization
von: Prabhune, Sonal, et al.
Veröffentlicht: (2025)
von: Prabhune, Sonal, et al.
Veröffentlicht: (2025)
Group Orthogonalized Policy Optimization:Group Policy Optimization as Orthogonal Projection in Hilbert Space
von: Zixian, Wang
Veröffentlicht: (2026)
von: Zixian, Wang
Veröffentlicht: (2026)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
von: Wang, Jialu, et al.
Veröffentlicht: (2026)
von: Wang, Jialu, et al.
Veröffentlicht: (2026)
Knowledgeable Agents by Offline Reinforcement Learning from Large Language Model Rollouts
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
von: Pang, Jing-Cheng, et al.
Veröffentlicht: (2024)
CoPRIS: Efficient and Stable Reinforcement Learning via Concurrency-Controlled Partial Rollout with Importance Sampling
von: Qu, Zekai, et al.
Veröffentlicht: (2025)
von: Qu, Zekai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GRAPH-GRPO-LEX: Contract Graph Modeling and Reinforcement Learning with Group Relative Policy Optimization
von: Dechtiar, Moriya, et al.
Veröffentlicht: (2025) -
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering
von: Cai, Yuzhu, et al.
Veröffentlicht: (2026) -
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025) -
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
von: Wang, Jingyi, et al.
Veröffentlicht: (2026) -
EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance
von: Yu, Song, et al.
Veröffentlicht: (2026)