Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
Fuente:
arXiv
Guardado en:
| Autores principales: | Roth, Amit, Samanta, Ankur, Halevy, Matan, Levine, Yoav, Efroni, Yonathan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reward Hacking Mitigation using Verifiable Composite Rewards
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
por: Helff, Lukas, et al.
Publicado: (2026)
por: Helff, Lukas, et al.
Publicado: (2026)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)
por: Ackermann, Johannes, et al.
Publicado: (2026)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
por: Beigi, Mohammad, et al.
Publicado: (2026)
por: Beigi, Mohammad, et al.
Publicado: (2026)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
por: Deshpande, Darshan, et al.
Publicado: (2026)
por: Deshpande, Darshan, et al.
Publicado: (2026)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
por: Wang, Chaoqi, et al.
Publicado: (2025)
por: Wang, Chaoqi, et al.
Publicado: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
Gradient Free Deep Reinforcement Learning With TabPFN
por: Schiff, David, et al.
Publicado: (2025)
por: Schiff, David, et al.
Publicado: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
por: Miao, Yuchun, et al.
Publicado: (2024)
por: Miao, Yuchun, et al.
Publicado: (2024)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
por: Thaman, Kunvar
Publicado: (2026)
por: Thaman, Kunvar
Publicado: (2026)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
por: Laidlaw, Cassidy, et al.
Publicado: (2024)
por: Laidlaw, Cassidy, et al.
Publicado: (2024)
Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do
por: Wald, Yoav, et al.
Publicado: (2025)
por: Wald, Yoav, et al.
Publicado: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
por: Pan, Alexander, et al.
Publicado: (2024)
por: Pan, Alexander, et al.
Publicado: (2024)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
por: Kwon, Jeongyeol, et al.
Publicado: (2024)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
por: Singha, Disha
Publicado: (2026)
por: Singha, Disha
Publicado: (2026)
HackAtari: Atari Learning Environments for Robust and Continual Reinforcement Learning
por: Delfosse, Quentin, et al.
Publicado: (2024)
por: Delfosse, Quentin, et al.
Publicado: (2024)
GARDO: Reinforcing Diffusion Models without Reward Hacking
por: He, Haoran, et al.
Publicado: (2025)
por: He, Haoran, et al.
Publicado: (2025)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
por: Beigi, Mohammad, et al.
Publicado: (2026)
por: Beigi, Mohammad, et al.
Publicado: (2026)
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
por: McKee-Reid, Leo, et al.
Publicado: (2024)
por: McKee-Reid, Leo, et al.
Publicado: (2024)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
por: Wu, Runzhe, et al.
Publicado: (2025)
por: Wu, Runzhe, et al.
Publicado: (2025)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
por: Farquhar, Sebastian, et al.
Publicado: (2025)
por: Farquhar, Sebastian, et al.
Publicado: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
por: Khalifa, Muhammad, et al.
Publicado: (2026)
por: Khalifa, Muhammad, et al.
Publicado: (2026)
Mitigating Preference Hacking in Policy Optimization with Pessimism
por: Gupta, Dhawal, et al.
Publicado: (2025)
por: Gupta, Dhawal, et al.
Publicado: (2025)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
por: Rashidinejad, Paria, et al.
Publicado: (2024)
por: Rashidinejad, Paria, et al.
Publicado: (2024)
On Teacher Hacking in Language Model Distillation
por: Tiapkin, Daniil, et al.
Publicado: (2025)
por: Tiapkin, Daniil, et al.
Publicado: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
por: Sahoo, Subramanyam
Publicado: (2026)
por: Sahoo, Subramanyam
Publicado: (2026)
Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
por: Wu, Lili, et al.
Publicado: (2024)
por: Wu, Lili, et al.
Publicado: (2024)
Fairness Hacking: The Malicious Practice of Shrouding Unfairness in Algorithms
por: Meding, Kristof, et al.
Publicado: (2023)
por: Meding, Kristof, et al.
Publicado: (2023)
LuckyMera: a Modular AI Framework for Building Hybrid NetHack Agents
por: Quarantiello, Luigi, et al.
Publicado: (2023)
por: Quarantiello, Luigi, et al.
Publicado: (2023)
School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
por: Taylor, Mia, et al.
Publicado: (2025)
por: Taylor, Mia, et al.
Publicado: (2025)
Defining and Characterizing Reward Hacking
por: Skalse, Joar, et al.
Publicado: (2022)
por: Skalse, Joar, et al.
Publicado: (2022)
Reward Hacking as Equilibrium under Finite Evaluation
por: Wang, Jiacheng, et al.
Publicado: (2026)
por: Wang, Jiacheng, et al.
Publicado: (2026)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
por: Baumann, Joachim, et al.
Publicado: (2025)
por: Baumann, Joachim, et al.
Publicado: (2025)
Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems
por: Olukola, Oluseyi, et al.
Publicado: (2026)
por: Olukola, Oluseyi, et al.
Publicado: (2026)
Reward Hacking in Rubric-Based Reinforcement Learning
por: Mahmoud, Anas, et al.
Publicado: (2026)
por: Mahmoud, Anas, et al.
Publicado: (2026)
Inferring Discussion Topics about Exploitation of Vulnerabilities from Underground Hacking Forums
por: Moreno-Vera, Felipe
Publicado: (2024)
por: Moreno-Vera, Felipe
Publicado: (2024)
Ejemplares similares
-
Reward Hacking Mitigation using Verifiable Composite Rewards
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025) -
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
por: Helff, Lukas, et al.
Publicado: (2026) -
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025) -
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025) -
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)