Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation
Fuente:
arXiv
Saved in:
| Main Authors: | He, Xixiang, Sun, Qiyao, Cheng, Ao, Li, Xingming, Ji, Xuanyu, Lu, Hailun, Huang, Runke, Hu, Qingyong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Hallucination Detection: Adaptive Bayesian Estimation of Semantic Entropy with Guided Semantic Exploration
by: Sun, Qiyao, et al.
Published: (2026)
by: Sun, Qiyao, et al.
Published: (2026)
ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
by: Cheng, Ao, et al.
Published: (2026)
by: Cheng, Ao, et al.
Published: (2026)
StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning
by: He, Xixiang, et al.
Published: (2026)
by: He, Xixiang, et al.
Published: (2026)
GAGPO: Generalized Advantage Grouped Policy Optimization
by: Zhu, Siyuan, et al.
Published: (2026)
by: Zhu, Siyuan, et al.
Published: (2026)
Your Group-Relative Advantage Is Biased
by: Yang, Fengkai, et al.
Published: (2026)
by: Yang, Fengkai, et al.
Published: (2026)
CAWR: Corruption-Averse Advantage-Weighted Regression for Robust Policy Optimization
by: Hu, Ranting
Published: (2025)
by: Hu, Ranting
Published: (2025)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
by: Ichihara, Yuki, et al.
Published: (2025)
by: Ichihara, Yuki, et al.
Published: (2025)
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization
by: Chu, Xu, et al.
Published: (2026)
by: Chu, Xu, et al.
Published: (2026)
NGRPO: Negative-enhanced Group Relative Policy Optimization
by: Nan, Gongrui, et al.
Published: (2025)
by: Nan, Gongrui, et al.
Published: (2025)
When AI Meets Early Childhood Education: Large Language Models as Assessment Teammates in Chinese Preschools
by: Li, Xingming, et al.
Published: (2026)
by: Li, Xingming, et al.
Published: (2026)
Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
by: Mroueh, Youssef, et al.
Published: (2025)
by: Mroueh, Youssef, et al.
Published: (2025)
Consensus Group Relative Policy Optimization for Text Generation
by: Ichihara, Yuki, et al.
Published: (2026)
by: Ichihara, Yuki, et al.
Published: (2026)
Reinforcement Unlearning via Group Relative Policy Optimization
by: Zaradoukas, Efstratios, et al.
Published: (2026)
by: Zaradoukas, Efstratios, et al.
Published: (2026)
Constrained Group Relative Policy Optimization
by: Girgis, Roger, et al.
Published: (2026)
by: Girgis, Roger, et al.
Published: (2026)
PepEVOLVE: Position-Aware Dynamic Peptide Optimization via Group-Relative Advantage
by: Nguyen, Trieu, et al.
Published: (2025)
by: Nguyen, Trieu, et al.
Published: (2025)
GRPOformer: Advancing Hyperparameter Optimization via Group Relative Policy Optimization
by: Guo, Haoxin, et al.
Published: (2025)
by: Guo, Haoxin, et al.
Published: (2025)
REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Research on Personalized Medical Intervention Strategy Generation System based on Group Relative Policy Optimization and Time-Series Data Fusion
by: Lu, Dingxin, et al.
Published: (2025)
by: Lu, Dingxin, et al.
Published: (2025)
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
Skip-Connected Policy Optimization for Implicit Advantage
by: Teng, Fengwei, et al.
Published: (2026)
by: Teng, Fengwei, et al.
Published: (2026)
Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic
by: Zhou, Hongyi, et al.
Published: (2026)
by: Zhou, Hongyi, et al.
Published: (2026)
The Value of Variance: Mitigating Debate Collapse in Multi-Agent Systems via Uncertainty-Driven Policy Optimization
by: Tang, Luoxi, et al.
Published: (2026)
by: Tang, Luoxi, et al.
Published: (2026)
Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
by: Le, Tue, et al.
Published: (2025)
by: Le, Tue, et al.
Published: (2025)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
by: Mundada, Gagan, et al.
Published: (2026)
by: Mundada, Gagan, et al.
Published: (2026)
Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
by: He, Shuo, et al.
Published: (2026)
by: He, Shuo, et al.
Published: (2026)
Hybrid Group Relative Policy Optimization: A Multi-Sample Approach to Enhancing Policy Optimization
by: Sane, Soham
Published: (2025)
by: Sane, Soham
Published: (2025)
Smooth Gate Functions for Soft Advantage Policy Optimization
by: Denisov, Egor, et al.
Published: (2026)
by: Denisov, Egor, et al.
Published: (2026)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
by: Kwon, Soo Min, et al.
Published: (2026)
by: Kwon, Soo Min, et al.
Published: (2026)
Policy Optimization via Adv2: Adversarial Learning on Advantage Functions
by: Jonckheere, Matthieu, et al.
Published: (2023)
by: Jonckheere, Matthieu, et al.
Published: (2023)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
by: Deng, Jingcheng, et al.
Published: (2026)
by: Deng, Jingcheng, et al.
Published: (2026)
An Advantage-based Optimization Method for Reinforcement Learning in Large Action Space
by: Lin, Hai, et al.
Published: (2024)
by: Lin, Hai, et al.
Published: (2024)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
Towards Flash Thinking via Decoupled Advantage Policy Optimization
by: Tan, Zezhong, et al.
Published: (2025)
by: Tan, Zezhong, et al.
Published: (2025)
Relative Policy-Transition Optimization for Fast Policy Transfer
by: Xu, Jiawei, et al.
Published: (2022)
by: Xu, Jiawei, et al.
Published: (2022)
F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
by: Surana, Rohan, et al.
Published: (2026)
by: Surana, Rohan, et al.
Published: (2026)
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering
by: Cai, Yuzhu, et al.
Published: (2026)
by: Cai, Yuzhu, et al.
Published: (2026)
Amortized Molecular Optimization via Group Relative Policy Optimization
by: Javaid, Muhammad bin, et al.
Published: (2026)
by: Javaid, Muhammad bin, et al.
Published: (2026)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
by: Wang, Jingyi, et al.
Published: (2026)
by: Wang, Jingyi, et al.
Published: (2026)
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
by: Wang, Jialu, et al.
Published: (2026)
by: Wang, Jialu, et al.
Published: (2026)
A Probabilistic Perspective on Model Collapse
by: Xu, Shirong, et al.
Published: (2025)
by: Xu, Shirong, et al.
Published: (2025)
Similar Items
-
Efficient Hallucination Detection: Adaptive Bayesian Estimation of Semantic Entropy with Guided Semantic Exploration
by: Sun, Qiyao, et al.
Published: (2026) -
ENC-Bench: A Benchmark for Evaluating Multimodal Large Language Models in Electronic Navigational Chart Understanding
by: Cheng, Ao, et al.
Published: (2026) -
StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning
by: He, Xixiang, et al.
Published: (2026) -
GAGPO: Generalized Advantage Grouped Policy Optimization
by: Zhu, Siyuan, et al.
Published: (2026) -
Your Group-Relative Advantage Is Biased
by: Yang, Fengkai, et al.
Published: (2026)