Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Zixuan, Sun, Xiaolin, Zheng, Zizhan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Enhancing LLM Safety via Constrained Direct Preference Optimization
di: Liu, Zixuan, et al.
Pubblicazione: (2024)
di: Liu, Zixuan, et al.
Pubblicazione: (2024)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2024)
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2024)
Belief-Enriched Pessimistic Q-Learning against Adversarial State Perturbations
di: Sun, Xiaolin, et al.
Pubblicazione: (2024)
di: Sun, Xiaolin, et al.
Pubblicazione: (2024)
Reward Hacking Mitigation using Verifiable Composite Rewards
di: Tarek, Mirza Farhan Bin, et al.
Pubblicazione: (2025)
di: Tarek, Mirza Farhan Bin, et al.
Pubblicazione: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025)
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
di: Eisenstein, Jacob, et al.
Pubblicazione: (2023)
di: Eisenstein, Jacob, et al.
Pubblicazione: (2023)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
di: Miao, Yuchun, et al.
Pubblicazione: (2025)
di: Miao, Yuchun, et al.
Pubblicazione: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
di: Chen, Lichang, et al.
Pubblicazione: (2024)
di: Chen, Lichang, et al.
Pubblicazione: (2024)
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
di: Song, Ruike, et al.
Pubblicazione: (2025)
di: Song, Ruike, et al.
Pubblicazione: (2025)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
di: Ichihara, Yuki, et al.
Pubblicazione: (2025)
di: Ichihara, Yuki, et al.
Pubblicazione: (2025)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
di: Yu, Zhuohao, et al.
Pubblicazione: (2026)
di: Yu, Zhuohao, et al.
Pubblicazione: (2026)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
di: Singha, Disha
Pubblicazione: (2026)
di: Singha, Disha
Pubblicazione: (2026)
UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
di: Fu, Lingling, et al.
Pubblicazione: (2025)
di: Fu, Lingling, et al.
Pubblicazione: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
di: Farquhar, Sebastian, et al.
Pubblicazione: (2025)
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
di: Miao, Yuchun, et al.
Pubblicazione: (2025)
di: Miao, Yuchun, et al.
Pubblicazione: (2025)
Defining and Characterizing Reward Hacking
di: Skalse, Joar, et al.
Pubblicazione: (2022)
di: Skalse, Joar, et al.
Pubblicazione: (2022)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
di: Wu, Rui, et al.
Pubblicazione: (2026)
di: Wu, Rui, et al.
Pubblicazione: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
di: Deng, Wenlong, et al.
Pubblicazione: (2026)
di: Deng, Wenlong, et al.
Pubblicazione: (2026)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
di: Lian, Jiesong, et al.
Pubblicazione: (2025)
di: Lian, Jiesong, et al.
Pubblicazione: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
di: Wang, Ye, et al.
Pubblicazione: (2026)
di: Wang, Ye, et al.
Pubblicazione: (2026)
Mitigating Preference Hacking in Policy Optimization with Pessimism
di: Gupta, Dhawal, et al.
Pubblicazione: (2025)
di: Gupta, Dhawal, et al.
Pubblicazione: (2025)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
di: Roth, Amit, et al.
Pubblicazione: (2026)
di: Roth, Amit, et al.
Pubblicazione: (2026)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
di: Zhai, Kevin, et al.
Pubblicazione: (2025)
di: Zhai, Kevin, et al.
Pubblicazione: (2025)
The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
di: Wu, Yusong, et al.
Pubblicazione: (2025)
di: Wu, Yusong, et al.
Pubblicazione: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
di: Wang, Chaoqi, et al.
Pubblicazione: (2025)
di: Wang, Chaoqi, et al.
Pubblicazione: (2025)
Inference-Time Reward Hacking in Large Language Models
di: Khalaf, Hadi, et al.
Pubblicazione: (2025)
di: Khalaf, Hadi, et al.
Pubblicazione: (2025)
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
di: Wang, Xiaohua, et al.
Pubblicazione: (2026)
di: Wang, Xiaohua, et al.
Pubblicazione: (2026)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
di: Rashidinejad, Paria, et al.
Pubblicazione: (2024)
di: Rashidinejad, Paria, et al.
Pubblicazione: (2024)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
di: Wang, Songtao, et al.
Pubblicazione: (2026)
di: Wang, Songtao, et al.
Pubblicazione: (2026)
EvilGenie: A Reward Hacking Benchmark
di: Gabor, Jonathan, et al.
Pubblicazione: (2025)
di: Gabor, Jonathan, et al.
Pubblicazione: (2025)
Do Synthetic Trajectories Reflect Real Reward Hacking? A Systematic Study on Monitoring In-the-Wild Hacking in Code Generation
di: Li, Lichen, et al.
Pubblicazione: (2026)
di: Li, Lichen, et al.
Pubblicazione: (2026)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
di: Helff, Lukas, et al.
Pubblicazione: (2026)
di: Helff, Lukas, et al.
Pubblicazione: (2026)
Insider Attacks in Multi-Agent LLM Consensus Systems
di: Sun, Xiaolin, et al.
Pubblicazione: (2026)
di: Sun, Xiaolin, et al.
Pubblicazione: (2026)
Fair Algorithms with Probing for Multi-Agent Multi-Armed Bandits
di: Xu, Tianyi, et al.
Pubblicazione: (2025)
di: Xu, Tianyi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Enhancing LLM Safety via Constrained Direct Preference Optimization
di: Liu, Zixuan, et al.
Pubblicazione: (2024) -
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
di: Laidlaw, Cassidy, et al.
Pubblicazione: (2024) -
Belief-Enriched Pessimistic Q-Learning against Adversarial State Perturbations
di: Sun, Xiaolin, et al.
Pubblicazione: (2024) -
Reward Hacking Mitigation using Verifiable Composite Rewards
di: Tarek, Mirza Farhan Bin, et al.
Pubblicazione: (2025) -
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025)