GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shih-Yang, Dong, Xin, Lu, Ximing, Diao, Shizhe, Belcak, Peter, Liu, Mingjie, Chen, Min-Hung, Yin, Hongxu, Wang, Yu-Chiang Frank, Cheng, Kwang-Ting, Choi, Yejin, Kautz, Jan, Molchanov, Pavlo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
by: Liu, Shih-Yang, et al.
Published: (2025)
by: Liu, Shih-Yang, et al.
Published: (2025)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization
by: Jin, Zhangyu, et al.
Published: (2026)
by: Jin, Zhangyu, et al.
Published: (2026)
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
by: Su, Hongjin, et al.
Published: (2025)
by: Su, Hongjin, et al.
Published: (2025)
Universal Deep Research: Bring Your Own Model and Strategy
by: Belcak, Peter, et al.
Published: (2025)
by: Belcak, Peter, et al.
Published: (2025)
Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation
by: Belcak, Peter, et al.
Published: (2025)
by: Belcak, Peter, et al.
Published: (2025)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
by: Kang, Minki, et al.
Published: (2026)
by: Kang, Minki, et al.
Published: (2026)
DoRA: Weight-Decomposed Low-Rank Adaptation
by: Liu, Shih-Yang, et al.
Published: (2024)
by: Liu, Shih-Yang, et al.
Published: (2024)
Small Language Models are the Future of Agentic AI
by: Belcak, Peter, et al.
Published: (2025)
by: Belcak, Peter, et al.
Published: (2025)
Semantic segmentation with reward
by: Ting, Xie, et al.
Published: (2025)
by: Ting, Xie, et al.
Published: (2025)
Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
by: Diao, Shizhe, et al.
Published: (2025)
by: Diao, Shizhe, et al.
Published: (2025)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
by: Zhang, Hao, et al.
Published: (2026)
by: Zhang, Hao, et al.
Published: (2026)
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
by: Hatamizadeh, Ali, et al.
Published: (2026)
by: Hatamizadeh, Ali, et al.
Published: (2026)
MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge through Group Direct Preference Optimization
by: Zhang, Lanxue, et al.
Published: (2025)
by: Zhang, Lanxue, et al.
Published: (2025)
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
by: Jiang, Guochao, et al.
Published: (2026)
by: Jiang, Guochao, et al.
Published: (2026)
Scaling RL to Long Videos
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
LITA: Language Instructed Temporal-Localization Assistant
by: Huang, De-An, et al.
Published: (2024)
by: Huang, De-An, et al.
Published: (2024)
Flextron: Many-in-One Flexible Large Language Model
by: Cai, Ruisi, et al.
Published: (2024)
by: Cai, Ruisi, et al.
Published: (2024)
REC-RL: Referring expression counting via Gaussian and range-based reward optimization
by: Liu, Hui, et al.
Published: (2026)
by: Liu, Hui, et al.
Published: (2026)
Autonomous state-space segmentation for Deep-RL sparse reward scenarios
by: Maselli, Gianluca, et al.
Published: (2025)
by: Maselli, Gianluca, et al.
Published: (2025)
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
by: Xu, Ruijie, et al.
Published: (2024)
by: Xu, Ruijie, et al.
Published: (2024)
EoRA: Fine-tuning-free Compensation for Compressed LLM with Eigenspace Low-Rank Approximation
by: Liu, Shih-Yang, et al.
Published: (2024)
by: Liu, Shih-Yang, et al.
Published: (2024)
AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One
by: Ranzinger, Mike, et al.
Published: (2023)
by: Ranzinger, Mike, et al.
Published: (2023)
FasterViT: Fast Vision Transformers with Hierarchical Attention
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
Polar: Agentic RL on Any Harness at Scale
by: Xu, Binfeng, et al.
Published: (2026)
by: Xu, Binfeng, et al.
Published: (2026)
TractOracle: towards an anatomically-informed reward function for RL-based tractography
by: Théberge, Antoine, et al.
Published: (2024)
by: Théberge, Antoine, et al.
Published: (2024)
Hymba: A Hybrid-head Architecture for Small Language Models
by: Dong, Xin, et al.
Published: (2024)
by: Dong, Xin, et al.
Published: (2024)
GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-Resolution
by: Yi, Qiaosi, et al.
Published: (2026)
by: Yi, Qiaosi, et al.
Published: (2026)
LIRE: listwise reward enhancement for preference alignment
by: Zhu, Mingye, et al.
Published: (2024)
by: Zhu, Mingye, et al.
Published: (2024)
Credit card rewards counterpoint
by: Aaron Klein, et al.
Published: (2025)
by: Aaron Klein, et al.
Published: (2025)
Science and technology. virtue rewarded
Published: (2002)
Published: (2002)
Discounting of shared rewards in pigeons
by: Tetsuo Yamaguchi
Published: (2019)
by: Tetsuo Yamaguchi
Published: (2019)
LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement
by: Ye, Zhifan, et al.
Published: (2025)
by: Ye, Zhifan, et al.
Published: (2025)
Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
by: Lu, Ximing, et al.
Published: (2026)
by: Lu, Ximing, et al.
Published: (2026)
Episodic Reinforcement Learning with Expanded State-reward Space
by: Liang, Dayang, et al.
Published: (2024)
by: Liang, Dayang, et al.
Published: (2024)
Adaptive Sharpness-Aware Pruning for Robust Sparse Networks
by: Bair, Anna, et al.
Published: (2023)
by: Bair, Anna, et al.
Published: (2023)
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
by: Heinrich, Greg, et al.
Published: (2024)
by: Heinrich, Greg, et al.
Published: (2024)
Similar Items
-
DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning
by: Liu, Shih-Yang, et al.
Published: (2025) -
BroRL: Scaling Reinforcement Learning via Broadened Exploration
by: Hu, Jian, et al.
Published: (2025) -
GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization
by: Jin, Zhangyu, et al.
Published: (2026) -
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
by: Liu, Mingjie, et al.
Published: (2025) -
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
by: Su, Hongjin, et al.
Published: (2025)