GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Shih-Yang, Dong, Xin, Lu, Ximing, Diao, Shizhe, Belcak, Peter, Liu, Mingjie, Chen, Min-Hung, Yin, Hongxu, Wang, Yu-Chiang Frank, Cheng, Kwang-Ting, Choi, Yejin, Kautz, Jan, Molchanov, Pavlo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917190931841024
author Liu, Shih-Yang
Dong, Xin
Lu, Ximing
Diao, Shizhe
Belcak, Peter
Liu, Mingjie
Chen, Min-Hung
Yin, Hongxu
Wang, Yu-Chiang Frank
Cheng, Kwang-Ting
Choi, Yejin
Kautz, Jan
Molchanov, Pavlo
author_facet Liu, Shih-Yang
Dong, Xin
Lu, Ximing
Diao, Shizhe
Belcak, Peter
Liu, Mingjie
Chen, Min-Hung
Yin, Hongxu
Wang, Yu-Chiang Frank
Cheng, Kwang-Ting
Choi, Yejin
Kautz, Jan
Molchanov, Pavlo
contents As language models become increasingly capable, users expect them to provide not only accurate responses but also behaviors aligned with diverse human preferences across a variety of scenarios. To achieve this, Reinforcement learning (RL) pipelines have begun incorporating multiple rewards, each capturing a distinct preference, to guide models toward these desired behaviors. However, recent work has defaulted to apply Group Relative Policy Optimization (GRPO) under multi-reward setting without examining its suitability. In this paper, we demonstrate that directly applying GRPO to normalize distinct rollout reward combinations causes them to collapse into identical advantage values, reducing the resolution of the training signal and resulting in suboptimal convergence and, in some cases, early training failure. We then introduce Group reward-Decoupled Normalization Policy Optimization (GDPO), a new policy optimization method to resolve these issues by decoupling the normalization of individual rewards, more faithfully preserving their relative differences and enabling more accurate multi-reward optimization, along with substantially improved training stability. We compare GDPO with GRPO across three tasks: tool calling, math reasoning, and coding reasoning, evaluating both correctness metrics (accuracy, bug ratio) and constraint adherence metrics (format, length). Across all settings, GDPO consistently outperforms GRPO, demonstrating its effectiveness and generalizability for multi-reward reinforcement learning optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2601_05242
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Liu, Shih-Yang
Dong, Xin
Lu, Ximing
Diao, Shizhe
Belcak, Peter
Liu, Mingjie
Chen, Min-Hung
Yin, Hongxu
Wang, Yu-Chiang Frank
Cheng, Kwang-Ting
Choi, Yejin
Kautz, Jan
Molchanov, Pavlo
Computation and Language
Artificial Intelligence
Machine Learning
As language models become increasingly capable, users expect them to provide not only accurate responses but also behaviors aligned with diverse human preferences across a variety of scenarios. To achieve this, Reinforcement learning (RL) pipelines have begun incorporating multiple rewards, each capturing a distinct preference, to guide models toward these desired behaviors. However, recent work has defaulted to apply Group Relative Policy Optimization (GRPO) under multi-reward setting without examining its suitability. In this paper, we demonstrate that directly applying GRPO to normalize distinct rollout reward combinations causes them to collapse into identical advantage values, reducing the resolution of the training signal and resulting in suboptimal convergence and, in some cases, early training failure. We then introduce Group reward-Decoupled Normalization Policy Optimization (GDPO), a new policy optimization method to resolve these issues by decoupling the normalization of individual rewards, more faithfully preserving their relative differences and enabling more accurate multi-reward optimization, along with substantially improved training stability. We compare GDPO with GRPO across three tasks: tool calling, math reasoning, and coding reasoning, evaluating both correctness metrics (accuracy, bug ratio) and constraint adherence metrics (format, length). Across all settings, GDPO consistently outperforms GRPO, demonstrating its effectiveness and generalizability for multi-reward reinforcement learning optimization.
title GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.05242