Do Synthetic Trajectories Reflect Real Reward Hacking? A Systematic Study on Monitoring In-the-Wild Hacking in Code Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Lichen, Zhou, Hengguang, Liang, Yijun, Zhou, Tianyi, Hsieh, Cho-Jui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding Reward Hacking in Text-to-Image Reinforcement Learning
por: Hong, Yunqi, et al.
Publicado: (2026)
por: Hong, Yunqi, et al.
Publicado: (2026)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
por: Roth, Amit, et al.
Publicado: (2026)
por: Roth, Amit, et al.
Publicado: (2026)
Defining and Characterizing Reward Hacking
por: Skalse, Joar, et al.
Publicado: (2022)
por: Skalse, Joar, et al.
Publicado: (2022)
R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
por: Zhou, Hengguang, et al.
Publicado: (2025)
por: Zhou, Hengguang, et al.
Publicado: (2025)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
por: Khalifa, Muhammad, et al.
Publicado: (2026)
por: Khalifa, Muhammad, et al.
Publicado: (2026)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
por: Lian, Jiesong, et al.
Publicado: (2025)
por: Lian, Jiesong, et al.
Publicado: (2025)
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
por: Li, Xirui, et al.
Publicado: (2024)
por: Li, Xirui, et al.
Publicado: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
por: Miao, Yuchun, et al.
Publicado: (2025)
por: Miao, Yuchun, et al.
Publicado: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
por: Wang, Songtao, et al.
Publicado: (2026)
por: Wang, Songtao, et al.
Publicado: (2026)
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
por: Liu, Zixuan, et al.
Publicado: (2026)
por: Liu, Zixuan, et al.
Publicado: (2026)
Inference-Time Reward Hacking in Large Language Models
por: Khalaf, Hadi, et al.
Publicado: (2025)
por: Khalaf, Hadi, et al.
Publicado: (2025)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
por: Deshpande, Darshan, et al.
Publicado: (2026)
por: Deshpande, Darshan, et al.
Publicado: (2026)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
por: Beigi, Mohammad, et al.
Publicado: (2026)
por: Beigi, Mohammad, et al.
Publicado: (2026)
EvilGenie: A Reward Hacking Benchmark
por: Gabor, Jonathan, et al.
Publicado: (2025)
por: Gabor, Jonathan, et al.
Publicado: (2025)
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
por: Wang, Xiaohua, et al.
Publicado: (2026)
por: Wang, Xiaohua, et al.
Publicado: (2026)
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
por: Miao, Yuchun, et al.
Publicado: (2025)
por: Miao, Yuchun, et al.
Publicado: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
por: Miao, Yuchun, et al.
Publicado: (2024)
por: Miao, Yuchun, et al.
Publicado: (2024)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
por: Eisenstein, Jacob, et al.
Publicado: (2023)
por: Eisenstein, Jacob, et al.
Publicado: (2023)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
por: Wang, Chaoqi, et al.
Publicado: (2025)
por: Wang, Chaoqi, et al.
Publicado: (2025)
GARDO: Reinforcing Diffusion Models without Reward Hacking
por: He, Haoran, et al.
Publicado: (2025)
por: He, Haoran, et al.
Publicado: (2025)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
por: Chen, Zihan, et al.
Publicado: (2025)
por: Chen, Zihan, et al.
Publicado: (2025)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
por: Singha, Disha
Publicado: (2026)
por: Singha, Disha
Publicado: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
por: Deng, Wenlong, et al.
Publicado: (2026)
por: Deng, Wenlong, et al.
Publicado: (2026)
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
por: Song, Ruike, et al.
Publicado: (2025)
por: Song, Ruike, et al.
Publicado: (2025)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
por: Yu, Zhuohao, et al.
Publicado: (2026)
por: Yu, Zhuohao, et al.
Publicado: (2026)
LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
por: Helff, Lukas, et al.
Publicado: (2026)
por: Helff, Lukas, et al.
Publicado: (2026)
Hacking Predictors Means Hacking Cars: Using Sensitivity Analysis to Identify Trajectory Prediction Vulnerabilities for Autonomous Driving Security
por: Gibson, Marsalis, et al.
Publicado: (2024)
por: Gibson, Marsalis, et al.
Publicado: (2024)
UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
por: Fu, Lingling, et al.
Publicado: (2025)
por: Fu, Lingling, et al.
Publicado: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
por: Pan, Alexander, et al.
Publicado: (2024)
por: Pan, Alexander, et al.
Publicado: (2024)
Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use
por: Thaman, Kunvar
Publicado: (2026)
por: Thaman, Kunvar
Publicado: (2026)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
por: Wu, Rui, et al.
Publicado: (2026)
por: Wu, Rui, et al.
Publicado: (2026)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
por: Laidlaw, Cassidy, et al.
Publicado: (2024)
por: Laidlaw, Cassidy, et al.
Publicado: (2024)
Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
por: Wu, Yusong, et al.
Publicado: (2025)
por: Wu, Yusong, et al.
Publicado: (2025)
Hacking Task Confounder in Meta-Learning
por: Wang, Jingyao, et al.
Publicado: (2023)
por: Wang, Jingyao, et al.
Publicado: (2023)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
por: Zhai, Kevin, et al.
Publicado: (2025)
por: Zhai, Kevin, et al.
Publicado: (2025)
Ejemplares similares
-
Understanding Reward Hacking in Text-to-Image Reinforcement Learning
por: Hong, Yunqi, et al.
Publicado: (2026) -
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024) -
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
por: Roth, Amit, et al.
Publicado: (2026) -
Defining and Characterizing Reward Hacking
por: Skalse, Joar, et al.
Publicado: (2022) -
R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
por: Zhou, Hengguang, et al.
Publicado: (2025)