A Unified Framework for Rethinking Policy Divergence Measures in GRPO
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Qingyuan, Wang, Yuhui, Zhan, Simon Sinong, Dai, Yanning, Deng, Shilong, Habchi, Sarra, Zhu, Qi, Gallé, Matthias, Huang, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Variational Delayed Policy Optimization
by: Wu, Qingyuan, et al.
Published: (2024)
by: Wu, Qingyuan, et al.
Published: (2024)
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
by: Zhan, Simon Sinong, et al.
Published: (2025)
by: Zhan, Simon Sinong, et al.
Published: (2025)
Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
by: Zhan, Simon Sinong, et al.
Published: (2024)
by: Zhan, Simon Sinong, et al.
Published: (2024)
Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays
by: Wu, Qingyuan, et al.
Published: (2024)
by: Wu, Qingyuan, et al.
Published: (2024)
Directly Forecasting Belief for Reinforcement Learning with Delays
by: Wu, Qingyuan, et al.
Published: (2025)
by: Wu, Qingyuan, et al.
Published: (2025)
MIMIC-Py: An Extensible Tool for Personality-Driven Automated Game Testing with Large Language Models
by: Chen, Yifei, et al.
Published: (2026)
by: Chen, Yifei, et al.
Published: (2026)
MIMIC: Integrating Diverse Personality Traits for Better Game Testing Using Large Language Model
by: Chen, Yifei, et al.
Published: (2025)
by: Chen, Yifei, et al.
Published: (2025)
An Empirical Study on Code Review Activity Prediction and Its Impact in Practice
by: Olewicki, Doriane, et al.
Published: (2024)
by: Olewicki, Doriane, et al.
Published: (2024)
Case Study: Runtime Safety Verification of Neural Network Controlled System
by: Yang, Frank, et al.
Published: (2024)
by: Yang, Frank, et al.
Published: (2024)
Inverse Delayed Reinforcement Learning
by: Zhan, Simon Sinong, et al.
Published: (2024)
by: Zhan, Simon Sinong, et al.
Published: (2024)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
Efficient Morphology-Control Co-Design via Stackelberg Proximal Policy Optimization
by: Dai, Yanning, et al.
Published: (2026)
by: Dai, Yanning, et al.
Published: (2026)
$λ$-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
by: Wang, Yining, et al.
Published: (2025)
by: Wang, Yining, et al.
Published: (2025)
OP-GRPO: Efficient Off-Policy GRPO for Flow-Matching Models
by: Zhang, Liyu, et al.
Published: (2026)
by: Zhang, Liyu, et al.
Published: (2026)
Empowering Autonomous Driving with Large Language Models: A Safety Perspective
by: Wang, Yixuan, et al.
Published: (2023)
by: Wang, Yixuan, et al.
Published: (2023)
On the Costs and Benefits of Adopting Lifelong Learning for Software Analytics -- Empirical Study on Brown Build and Risk Prediction
by: Olewicki, Doriane, et al.
Published: (2023)
by: Olewicki, Doriane, et al.
Published: (2023)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
by: Deng, Jingcheng, et al.
Published: (2026)
by: Deng, Jingcheng, et al.
Published: (2026)
F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
by: Surana, Rohan, et al.
Published: (2026)
by: Surana, Rohan, et al.
Published: (2026)
Unified Framework for Calculating Convex Roof Resource Measures
by: Zhu, Xuanran, et al.
Published: (2024)
by: Zhu, Xuanran, et al.
Published: (2024)
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025)
by: Xue, Zeyue, et al.
Published: (2025)
Shedding Light on VLN Robustness: A Black-box Framework for Indoor Lighting-based Adversarial Attack
by: Li, Chenyang, et al.
Published: (2025)
by: Li, Chenyang, et al.
Published: (2025)
Switching Controller Synthesis for Hybrid Systems Against STL Formulas
by: Su, Han, et al.
Published: (2024)
by: Su, Han, et al.
Published: (2024)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
by: Li, Yuming, et al.
Published: (2025)
by: Li, Yuming, et al.
Published: (2025)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
by: Chen, Minghan, et al.
Published: (2025)
by: Chen, Minghan, et al.
Published: (2025)
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
by: Wang, Jiaqi, et al.
Published: (2026)
by: Wang, Jiaqi, et al.
Published: (2026)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
by: Li, Yuhui, et al.
Published: (2024)
by: Li, Yuhui, et al.
Published: (2024)
How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learning
by: Tian, Minghao, et al.
Published: (2026)
by: Tian, Minghao, et al.
Published: (2026)
Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing
by: Xu, Shaodong, et al.
Published: (2026)
by: Xu, Shaodong, et al.
Published: (2026)
Unterrichtszentrierte Schulentwicklung
by: Galle, Marco
Published: (2021)
by: Galle, Marco
Published: (2021)
Geschichtsdarstellung in der Gegenwartsliteratur: Florian Illies‘ Pop-Chronik der Welt von Gestern
by: Helmut Galle
Published: (2014)
by: Helmut Galle
Published: (2014)
Zur Erinnerung an deutsche Opfer: Geschichte, Zeugnis und Fiktion in Grass’ Novelle Im Krebsgang.
by: Helmut Galle
Published: (2005)
by: Helmut Galle
Published: (2005)
Entre vítima e perpetrador: a identidade problemática da segunda geração pós- Shoá na Alemanha e a proposta do romance O leitor, de Bernhard Schlink
by: Helmut Galle
Published: (2007)
by: Helmut Galle
Published: (2007)
A volta da violência na política alemã. Estratégias legitimadoras na autobiografia de uma protagonista dos anos 1970
by: Helmut Galle
Published: (2009)
by: Helmut Galle
Published: (2009)
A poeta das “moradas da morte”. Sobre a obra lírica de Nelly Sachs
by: Helmut Galle
Published: (2006)
by: Helmut Galle
Published: (2006)
Einleitende Bemerkungen zum Thema des Dossiers: „Die deutsche Literatur und der Nobelpreis”
by: Helmut Galle
Published: (2006)
by: Helmut Galle
Published: (2006)
LLMCRIT: Teaching Large Language Models to Use Criteria
by: Yuan, Weizhe, et al.
Published: (2024)
by: Yuan, Weizhe, et al.
Published: (2024)
Scaling Value Iteration Networks to 5000 Layers for Extreme Long-Term Planning
by: Wang, Yuhui, et al.
Published: (2024)
by: Wang, Yuhui, et al.
Published: (2024)
Similar Items
-
Variational Delayed Policy Optimization
by: Wu, Qingyuan, et al.
Published: (2024) -
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
by: Zhan, Simon Sinong, et al.
Published: (2025) -
Enhancing Inverse Reinforcement Learning through Encoding Dynamic Information in Reward Shaping
by: Zhan, Simon Sinong, et al.
Published: (2024) -
Boosting Reinforcement Learning with Strongly Delayed Feedback Through Auxiliary Short Delays
by: Wu, Qingyuan, et al.
Published: (2024) -
Directly Forecasting Belief for Reinforcement Learning with Delays
by: Wu, Qingyuan, et al.
Published: (2025)