Mitigating Preference Hacking in Policy Optimization with Pessimism
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Dhawal, Fisch, Adam, Dann, Christoph, Agarwal, Alekh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Design Considerations in Offline Preference-based RL
by: Agarwal, Alekh, et al.
Published: (2025)
by: Agarwal, Alekh, et al.
Published: (2025)
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
by: Agarwal, Alekh, et al.
Published: (2022)
by: Agarwal, Alekh, et al.
Published: (2022)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
by: Yu, Zhuohao, et al.
Published: (2026)
by: Yu, Zhuohao, et al.
Published: (2026)
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
by: Huang, Jiawei, et al.
Published: (2025)
by: Huang, Jiawei, et al.
Published: (2025)
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
by: Aouali, Imad, et al.
Published: (2024)
by: Aouali, Imad, et al.
Published: (2024)
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
by: Eisenstein, Jacob, et al.
Published: (2023)
by: Eisenstein, Jacob, et al.
Published: (2023)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
by: Farquhar, Sebastian, et al.
Published: (2025)
by: Farquhar, Sebastian, et al.
Published: (2025)
Data-Driven Online Model Selection With Regret Guarantees
by: Pacchiano, Aldo, et al.
Published: (2023)
by: Pacchiano, Aldo, et al.
Published: (2023)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
by: Hsu, Hsiang, et al.
Published: (2026)
by: Hsu, Hsiang, et al.
Published: (2026)
Robust Preference Optimization through Reward Model Distillation
by: Fisch, Adam, et al.
Published: (2024)
by: Fisch, Adam, et al.
Published: (2024)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
by: Laidlaw, Cassidy, et al.
Published: (2024)
by: Laidlaw, Cassidy, et al.
Published: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
by: Wang, Kaiwen, et al.
Published: (2024)
by: Wang, Kaiwen, et al.
Published: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
Mitigating Mismatch within Reference-based Preference Optimization
by: Yuan, Suqin, et al.
Published: (2026)
by: Yuan, Suqin, et al.
Published: (2026)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
by: Miao, Yuchun, et al.
Published: (2024)
by: Miao, Yuchun, et al.
Published: (2024)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)
by: Roth, Amit, et al.
Published: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
by: Nishimori, Soichiro, et al.
Published: (2025)
by: Nishimori, Soichiro, et al.
Published: (2025)
Bootstrapping LLMs via Preference-Based Policy Optimization
by: Jia, Chen
Published: (2025)
by: Jia, Chen
Published: (2025)
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
by: Swamy, Gokul, et al.
Published: (2024)
by: Swamy, Gokul, et al.
Published: (2024)
Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning
by: Kang, Hyungkyu, et al.
Published: (2025)
by: Kang, Hyungkyu, et al.
Published: (2025)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
GOPO: Policy Optimization using Ranked Rewards
by: Choi, Kyuseong, et al.
Published: (2026)
by: Choi, Kyuseong, et al.
Published: (2026)
Harnessing Network Effect for Fake News Mitigation: Selecting Debunkers via Self-Imitation Learning
by: Xu, Xiaofei, et al.
Published: (2024)
by: Xu, Xiaofei, et al.
Published: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training
by: Moya, Christian, et al.
Published: (2026)
by: Moya, Christian, et al.
Published: (2026)
Cost-Optimal Active AI Model Evaluation
by: Angelopoulos, Anastasios N., et al.
Published: (2025)
by: Angelopoulos, Anastasios N., et al.
Published: (2025)
Evolutionary Policy Optimization
by: Wang, Jianren, et al.
Published: (2025)
by: Wang, Jianren, et al.
Published: (2025)
Policy-labeled Preference Learning: Is Preference Enough for RLHF?
by: Cho, Taehyun, et al.
Published: (2025)
by: Cho, Taehyun, et al.
Published: (2025)
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
by: Xu, Wenzhe, et al.
Published: (2026)
by: Xu, Wenzhe, et al.
Published: (2026)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
by: McKee-Reid, Leo, et al.
Published: (2024)
by: McKee-Reid, Leo, et al.
Published: (2024)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025)
by: Gupta, Taneesh, et al.
Published: (2025)
Rich Insights from Cheap Signals: Efficient Evaluations via Tensor Factorization
by: Polo, Felipe Maia, et al.
Published: (2026)
by: Polo, Felipe Maia, et al.
Published: (2026)
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
by: Zhao, Zihui, et al.
Published: (2025)
by: Zhao, Zihui, et al.
Published: (2025)
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
by: Ambadkar, Tanmay, et al.
Published: (2026)
by: Ambadkar, Tanmay, et al.
Published: (2026)
Similar Items
-
Design Considerations in Offline Preference-based RL
by: Agarwal, Alekh, et al.
Published: (2025) -
Non-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling
by: Agarwal, Alekh, et al.
Published: (2022) -
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
by: Yu, Zhuohao, et al.
Published: (2026) -
Can RLHF be More Efficient with Imperfect Reward Models? A Policy Coverage Perspective
by: Huang, Jiawei, et al.
Published: (2025) -
Unified PAC-Bayesian Study of Pessimism for Offline Policy Learning with Regularized Importance Sampling
by: Aouali, Imad, et al.
Published: (2024)