MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
Fuente:
arXiv
Saved in:
| Main Authors: | Ichihara, Yuki, Jinnai, Yuu, Morimura, Tetsuro, Sakamoto, Mitsuki, Mitsuhashi, Ryota, Uchibe, Eiji |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Consensus Group Relative Policy Optimization for Text Generation
by: Ichihara, Yuki, et al.
Published: (2026)
by: Ichihara, Yuki, et al.
Published: (2026)
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
by: Ichihara, Yuki, et al.
Published: (2025)
by: Ichihara, Yuki, et al.
Published: (2025)
Theoretical Guarantees for Minimum Bayes Risk Decoding
by: Ichihara, Yuki, et al.
Published: (2025)
by: Ichihara, Yuki, et al.
Published: (2025)
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024)
by: Morimura, Tetsuro, et al.
Published: (2024)
Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
On the True Distribution Approximation of Minimum Bayes-Risk Decoding
by: Ohashi, Atsumoto, et al.
Published: (2024)
by: Ohashi, Atsumoto, et al.
Published: (2024)
Generating Diverse and High-Quality Texts by Minimum Bayes Risk Decoding
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
Re-evaluating Minimum Bayes Risk Decoding for Automatic Speech Recognition
by: Jinnai, Yuu
Published: (2025)
by: Jinnai, Yuu
Published: (2025)
Does Cross-Cultural Alignment Change the Commonsense Morality of Language Models?
by: Jinnai, Yuu
Published: (2024)
by: Jinnai, Yuu
Published: (2024)
Reward-Punishment Reinforcement Learning with Maximum Entropy
by: Wang, Jiexin, et al.
Published: (2024)
by: Wang, Jiexin, et al.
Published: (2024)
Model-Based Minimum Bayes Risk Decoding for Text Generation
by: Jinnai, Yuu, et al.
Published: (2023)
by: Jinnai, Yuu, et al.
Published: (2023)
Policy Gradient with Kernel Quadrature
by: Hayakawa, Satoshi, et al.
Published: (2023)
by: Hayakawa, Satoshi, et al.
Published: (2023)
Annotation-Efficient Language Model Alignment via Diverse and Representative Response Texts
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport
by: Jinnai, Yuu
Published: (2025)
by: Jinnai, Yuu
Published: (2025)
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
by: Liu, Zixuan, et al.
Published: (2026)
by: Liu, Zixuan, et al.
Published: (2026)
Interaction Locality in Hierarchical Recursive Reasoning
by: Miyanishi, Yosuke, et al.
Published: (2026)
by: Miyanishi, Yosuke, et al.
Published: (2026)
Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
by: Deng, Jingcheng, et al.
Published: (2026)
by: Deng, Jingcheng, et al.
Published: (2026)
Mitigating Preference Hacking in Policy Optimization with Pessimism
by: Gupta, Dhawal, et al.
Published: (2025)
by: Gupta, Dhawal, et al.
Published: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
FedGRPO: Privately Optimizing Foundation Models with Group-Relative Rewards from Domain Client
by: Zhu, Gongxi, et al.
Published: (2026)
by: Zhu, Gongxi, et al.
Published: (2026)
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Policy Gradient Algorithms with Monte Carlo Tree Learning for Non-Markov Decision Processes
by: Morimura, Tetsuro, et al.
Published: (2022)
by: Morimura, Tetsuro, et al.
Published: (2022)
Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales
by: Tsutsumi, Ayuto, et al.
Published: (2025)
by: Tsutsumi, Ayuto, et al.
Published: (2025)
Hyperparameter-Free Approach for Faster Minimum Bayes Risk Decoding
by: Jinnai, Yuu, et al.
Published: (2024)
by: Jinnai, Yuu, et al.
Published: (2024)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
by: Mundada, Gagan, et al.
Published: (2026)
by: Mundada, Gagan, et al.
Published: (2026)
F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking
by: Surana, Rohan, et al.
Published: (2026)
by: Surana, Rohan, et al.
Published: (2026)
IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization
by: Wang, Shuai, et al.
Published: (2026)
by: Wang, Shuai, et al.
Published: (2026)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
by: Farquhar, Sebastian, et al.
Published: (2025)
by: Farquhar, Sebastian, et al.
Published: (2025)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
by: Zhang, Xichen, et al.
Published: (2025)
by: Zhang, Xichen, et al.
Published: (2025)
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model
by: Zhou, Renping, et al.
Published: (2025)
by: Zhou, Renping, et al.
Published: (2025)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization
by: Kwon, Soo Min, et al.
Published: (2026)
by: Kwon, Soo Min, et al.
Published: (2026)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
by: Eisenstein, Jacob, et al.
Published: (2023)
by: Eisenstein, Jacob, et al.
Published: (2023)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
by: Miao, Yuchun, et al.
Published: (2025)
by: Miao, Yuchun, et al.
Published: (2025)
MO-CAPO: Multi-Objective Cost-Aware Prompt Optimization
by: Büssing, Jan, et al.
Published: (2026)
by: Büssing, Jan, et al.
Published: (2026)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
by: Wang, Jingyi, et al.
Published: (2026)
by: Wang, Jingyi, et al.
Published: (2026)
MC-GRPO: Median-Centered Group Relative Policy Optimization for Small-Rollout Reinforcement Learning
by: Kim, Youngeun
Published: (2026)
by: Kim, Youngeun
Published: (2026)
Similar Items
-
Consensus Group Relative Policy Optimization for Text Generation
by: Ichihara, Yuki, et al.
Published: (2026) -
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
by: Ichihara, Yuki, et al.
Published: (2025) -
Theoretical Guarantees for Minimum Bayes Risk Decoding
by: Ichihara, Yuki, et al.
Published: (2025) -
Filtered Direct Preference Optimization
by: Morimura, Tetsuro, et al.
Published: (2024) -
Regularized Best-of-N Sampling with Minimum Bayes Risk Objective for Language Model Alignment
by: Jinnai, Yuu, et al.
Published: (2024)