MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
Fuente:
arXiv
Saved in:
| Main Authors: | Farquhar, Sebastian, Varma, Vikrant, Lindner, David, Elson, David, Biddulph, Caleb, Goodfellow, Ian, Shah, Rohin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extending MONA in Camera Dropbox: Reproduction, Learned Approval, and Design Implications for Reward-Hacking Mitigation
by: Heath, Nathan
Published: (2026)
by: Heath, Nathan
Published: (2026)
Realistic honeypot evaluations for scheming propensity
by: Krakovna, Victoria, et al.
Published: (2026)
by: Krakovna, Victoria, et al.
Published: (2026)
Quantifying the Necessity of Chain of Thought through Opaque Serial Depth
by: Brown-Cohen, Jonah, et al.
Published: (2026)
by: Brown-Cohen, Jonah, et al.
Published: (2026)
A Pragmatic Way to Measure Chain-of-Thought Monitorability
by: Emmons, Scott, et al.
Published: (2025)
by: Emmons, Scott, et al.
Published: (2025)
Gram: Assessing sabotage propensities via automated alignment auditing
by: Lindner, David, et al.
Published: (2026)
by: Lindner, David, et al.
Published: (2026)
Differentiating Policies for Non-Myopic Bayesian Optimization
by: Nwankwo, Darian, et al.
Published: (2024)
by: Nwankwo, Darian, et al.
Published: (2024)
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
by: Kaufmann, Max, et al.
Published: (2026)
by: Kaufmann, Max, et al.
Published: (2026)
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
by: Liu, Zixuan, et al.
Published: (2026)
by: Liu, Zixuan, et al.
Published: (2026)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
Non-Myopic Multi-Objective Bayesian Optimization
by: Belakaria, Syrine, et al.
Published: (2024)
by: Belakaria, Syrine, et al.
Published: (2024)
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
Consistency Training Helps Stop Sycophancy and Jailbreaks
by: Irpan, Alex, et al.
Published: (2025)
by: Irpan, Alex, et al.
Published: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
RRM: Robust Reward Model Training Mitigates Reward Hacking
by: Liu, Tianqi, et al.
Published: (2024)
by: Liu, Tianqi, et al.
Published: (2024)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
by: Ichihara, Yuki, et al.
Published: (2025)
by: Ichihara, Yuki, et al.
Published: (2025)
Non-Myopic Multifidelity Bayesian Optimization
by: Di Fiore, Francesco, et al.
Published: (2022)
by: Di Fiore, Francesco, et al.
Published: (2022)
Defining and Characterizing Reward Hacking
by: Skalse, Joar, et al.
Published: (2022)
by: Skalse, Joar, et al.
Published: (2022)
Improving Dictionary Learning with Gated Sparse Autoencoders
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
by: Rajamanoharan, Senthooran, et al.
Published: (2024)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
by: Jang, Eyon, et al.
Published: (2026)
by: Jang, Eyon, et al.
Published: (2026)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
by: Eisenstein, Jacob, et al.
Published: (2023)
by: Eisenstein, Jacob, et al.
Published: (2023)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
by: Miao, Yuchun, et al.
Published: (2025)
by: Miao, Yuchun, et al.
Published: (2025)
Mitigating Information Asymmetry in Two-Stage Contracts with Non-Myopic Agents
by: Dahleh, Munther A., et al.
Published: (2024)
by: Dahleh, Munther A., et al.
Published: (2024)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
by: Lian, Jiesong, et al.
Published: (2025)
by: Lian, Jiesong, et al.
Published: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
by: Zhai, Kevin, et al.
Published: (2025)
by: Zhai, Kevin, et al.
Published: (2025)
UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
by: Fu, Lingling, et al.
Published: (2025)
by: Fu, Lingling, et al.
Published: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
by: Miao, Yuchun, et al.
Published: (2024)
by: Miao, Yuchun, et al.
Published: (2024)
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
by: Song, Ruike, et al.
Published: (2025)
by: Song, Ruike, et al.
Published: (2025)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
by: Wu, Rui, et al.
Published: (2026)
by: Wu, Rui, et al.
Published: (2026)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
by: Yu, Zhuohao, et al.
Published: (2026)
by: Yu, Zhuohao, et al.
Published: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
by: Deng, Wenlong, et al.
Published: (2026)
by: Deng, Wenlong, et al.
Published: (2026)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
by: Laidlaw, Cassidy, et al.
Published: (2024)
by: Laidlaw, Cassidy, et al.
Published: (2024)
Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation
by: Liu, Yifeng, et al.
Published: (2026)
by: Liu, Yifeng, et al.
Published: (2026)
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
by: Ferreira, Pedro, et al.
Published: (2025)
by: Ferreira, Pedro, et al.
Published: (2025)
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
by: Li, Jiacheng, et al.
Published: (2026)
by: Li, Jiacheng, et al.
Published: (2026)
Mitigating Preference Hacking in Policy Optimization with Pessimism
by: Gupta, Dhawal, et al.
Published: (2025)
by: Gupta, Dhawal, et al.
Published: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Similar Items
-
Extending MONA in Camera Dropbox: Reproduction, Learned Approval, and Design Implications for Reward-Hacking Mitigation
by: Heath, Nathan
Published: (2026) -
Realistic honeypot evaluations for scheming propensity
by: Krakovna, Victoria, et al.
Published: (2026) -
Quantifying the Necessity of Chain of Thought through Opaque Serial Depth
by: Brown-Cohen, Jonah, et al.
Published: (2026) -
A Pragmatic Way to Measure Chain-of-Thought Monitorability
by: Emmons, Scott, et al.
Published: (2025) -
Gram: Assessing sabotage propensities via automated alignment auditing
by: Lindner, David, et al.
Published: (2026)