Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Fuente:
arXiv
Guardado en:
| Autores principales: | Laidlaw, Cassidy, Singhal, Shivam, Dragan, Anca |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
por: Laidlaw, Cassidy, et al.
Publicado: (2023)
por: Laidlaw, Cassidy, et al.
Publicado: (2023)
Bridging RL Theory and Practice with the Effective Horizon
por: Laidlaw, Cassidy, et al.
Publicado: (2023)
por: Laidlaw, Cassidy, et al.
Publicado: (2023)
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
por: Feng, Dylan, et al.
Publicado: (2026)
por: Feng, Dylan, et al.
Publicado: (2026)
AssistanceZero: Scalably Solving Assistance Games
por: Laidlaw, Cassidy, et al.
Publicado: (2025)
por: Laidlaw, Cassidy, et al.
Publicado: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
por: Beigi, Mohammad, et al.
Publicado: (2026)
por: Beigi, Mohammad, et al.
Publicado: (2026)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
por: Liu, Zixuan, et al.
Publicado: (2026)
por: Liu, Zixuan, et al.
Publicado: (2026)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
por: Siththaranjan, Anand, et al.
Publicado: (2023)
por: Siththaranjan, Anand, et al.
Publicado: (2023)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
AI Alignment with Changing and Influenceable Reward Functions
por: Carroll, Micah, et al.
Publicado: (2024)
por: Carroll, Micah, et al.
Publicado: (2024)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
por: Miao, Yuchun, et al.
Publicado: (2024)
por: Miao, Yuchun, et al.
Publicado: (2024)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
por: Ye, Yaowen, et al.
Publicado: (2025)
por: Ye, Yaowen, et al.
Publicado: (2025)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
por: Roth, Amit, et al.
Publicado: (2026)
por: Roth, Amit, et al.
Publicado: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
por: Beigi, Mohammad, et al.
Publicado: (2026)
por: Beigi, Mohammad, et al.
Publicado: (2026)
A Generalized Acquisition Function for Preference-based Reward Learning
por: Ellis, Evan, et al.
Publicado: (2024)
por: Ellis, Evan, et al.
Publicado: (2024)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
por: Singha, Disha
Publicado: (2026)
por: Singha, Disha
Publicado: (2026)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
por: Farquhar, Sebastian, et al.
Publicado: (2025)
por: Farquhar, Sebastian, et al.
Publicado: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
por: Wang, Chaoqi, et al.
Publicado: (2025)
por: Wang, Chaoqi, et al.
Publicado: (2025)
Mitigating Preference Hacking in Policy Optimization with Pessimism
por: Gupta, Dhawal, et al.
Publicado: (2025)
por: Gupta, Dhawal, et al.
Publicado: (2025)
Learning to Assist Humans without Inferring Rewards
por: Myers, Vivek, et al.
Publicado: (2024)
por: Myers, Vivek, et al.
Publicado: (2024)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
por: Helff, Lukas, et al.
Publicado: (2026)
por: Helff, Lukas, et al.
Publicado: (2026)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
por: Hong, Joey, et al.
Publicado: (2024)
por: Hong, Joey, et al.
Publicado: (2024)
Adversaries Can Misuse Combinations of Safe Models
por: Jones, Erik, et al.
Publicado: (2024)
por: Jones, Erik, et al.
Publicado: (2024)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
por: Thaman, Kunvar
Publicado: (2026)
por: Thaman, Kunvar
Publicado: (2026)
Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making
por: Myers, Vivek, et al.
Publicado: (2024)
por: Myers, Vivek, et al.
Publicado: (2024)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
por: McKee-Reid, Leo, et al.
Publicado: (2024)
por: McKee-Reid, Leo, et al.
Publicado: (2024)
Feedback Loops With Language Models Drive In-Context Reward Hacking
por: Pan, Alexander, et al.
Publicado: (2024)
por: Pan, Alexander, et al.
Publicado: (2024)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
por: Hong, Joey, et al.
Publicado: (2024)
por: Hong, Joey, et al.
Publicado: (2024)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)
por: Ackermann, Johannes, et al.
Publicado: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
por: Khalifa, Muhammad, et al.
Publicado: (2026)
por: Khalifa, Muhammad, et al.
Publicado: (2026)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
por: Deshpande, Darshan, et al.
Publicado: (2026)
por: Deshpande, Darshan, et al.
Publicado: (2026)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
por: Lang, Leon, et al.
Publicado: (2024)
por: Lang, Leon, et al.
Publicado: (2024)
On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback
por: Williams, Marcus, et al.
Publicado: (2024)
por: Williams, Marcus, et al.
Publicado: (2024)
Training LLM Agents to Empower Humans
por: Ellis, Evan, et al.
Publicado: (2025)
por: Ellis, Evan, et al.
Publicado: (2025)
GARDO: Reinforcing Diffusion Models without Reward Hacking
por: He, Haoran, et al.
Publicado: (2025)
por: He, Haoran, et al.
Publicado: (2025)
Rethinking the Role of Proxy Rewards in Language Model Alignment
por: Kim, Sungdong, et al.
Publicado: (2024)
por: Kim, Sungdong, et al.
Publicado: (2024)
Ejemplares similares
-
The Effective Horizon Explains Deep RL Performance in Stochastic Environments
por: Laidlaw, Cassidy, et al.
Publicado: (2023) -
Bridging RL Theory and Practice with the Effective Horizon
por: Laidlaw, Cassidy, et al.
Publicado: (2023) -
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs
por: Feng, Dylan, et al.
Publicado: (2026) -
AssistanceZero: Scalably Solving Assistance Games
por: Laidlaw, Cassidy, et al.
Publicado: (2025) -
Reward Hacking Mitigation using Verifiable Composite Rewards
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)