Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Ruike, Song, Zeen, Guo, Huijie, Qiang, Wenwen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reward Model Generalization for Compute-Aware Test-Time Reasoning
by: Song, Zeen, et al.
Published: (2025)
by: Song, Zeen, et al.
Published: (2025)
Hacking Task Confounder in Meta-Learning
by: Wang, Jingyao, et al.
Published: (2023)
by: Wang, Jingyao, et al.
Published: (2023)
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025)
by: Fu, Jiayi, et al.
Published: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
Learning to Reason without External Rewards
by: Zhao, Xuandong, et al.
Published: (2025)
by: Zhao, Xuandong, et al.
Published: (2025)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
by: Miao, Yuchun, et al.
Published: (2025)
by: Miao, Yuchun, et al.
Published: (2025)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
by: Eisenstein, Jacob, et al.
Published: (2023)
by: Eisenstein, Jacob, et al.
Published: (2023)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
by: Miao, Yuchun, et al.
Published: (2024)
by: Miao, Yuchun, et al.
Published: (2024)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
Not All Frequencies Are Created Equal:Towards a Dynamic Fusion of Frequencies in Time-Series Forecasting
by: Zhang, Xingyu, et al.
Published: (2024)
by: Zhang, Xingyu, et al.
Published: (2024)
Adaptive Uncertainty-Aware Tree Search for Robust Reasoning
by: Song, Zeen, et al.
Published: (2026)
by: Song, Zeen, et al.
Published: (2026)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
by: Singha, Disha
Published: (2026)
by: Singha, Disha
Published: (2026)
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
by: Liu, Zixuan, et al.
Published: (2026)
by: Liu, Zixuan, et al.
Published: (2026)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
by: Chen, Lichang, et al.
Published: (2024)
by: Chen, Lichang, et al.
Published: (2024)
From Shallow to Deep: Pinning Semantic Intent via Causal GRPO
by: Zhou, Shuyi, et al.
Published: (2026)
by: Zhou, Shuyi, et al.
Published: (2026)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
by: Lian, Jiesong, et al.
Published: (2025)
by: Lian, Jiesong, et al.
Published: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
by: Ono, Shinnosuke, et al.
Published: (2026)
by: Ono, Shinnosuke, et al.
Published: (2026)
UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
by: Fu, Lingling, et al.
Published: (2025)
by: Fu, Lingling, et al.
Published: (2025)
Beyond All-to-All: Causal-Aligned Transformer with Dynamic Structure Learning for Multivariate Time Series Forecasting
by: Zhang, Xingyu, et al.
Published: (2025)
by: Zhang, Xingyu, et al.
Published: (2025)
Defining and Characterizing Reward Hacking
by: Skalse, Joar, et al.
Published: (2022)
by: Skalse, Joar, et al.
Published: (2022)
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning
by: Wang, Jingyao, et al.
Published: (2026)
by: Wang, Jingyao, et al.
Published: (2026)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
by: Yu, Zhuohao, et al.
Published: (2026)
by: Yu, Zhuohao, et al.
Published: (2026)
Group Causal Policy Optimization for Post-Training Large Language Models
by: Gu, Ziyin, et al.
Published: (2025)
by: Gu, Ziyin, et al.
Published: (2025)
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
by: Miao, Yuchun, et al.
Published: (2025)
by: Miao, Yuchun, et al.
Published: (2025)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
by: Wu, Rui, et al.
Published: (2026)
by: Wu, Rui, et al.
Published: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
by: Deng, Wenlong, et al.
Published: (2026)
by: Deng, Wenlong, et al.
Published: (2026)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
by: Laidlaw, Cassidy, et al.
Published: (2024)
by: Laidlaw, Cassidy, et al.
Published: (2024)
On the Out-of-Distribution Generalization of Self-Supervised Learning
by: Qiang, Wenwen, et al.
Published: (2025)
by: Qiang, Wenwen, et al.
Published: (2025)
Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
by: Wang, Jingyao, et al.
Published: (2025)
by: Wang, Jingyao, et al.
Published: (2025)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
by: Roth, Amit, et al.
Published: (2026)
by: Roth, Amit, et al.
Published: (2026)
On the Generalization and Causal Explanation in Self-Supervised Learning
by: Qiang, Wenwen, et al.
Published: (2024)
by: Qiang, Wenwen, et al.
Published: (2024)
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
by: Zhai, Kevin, et al.
Published: (2025)
by: Zhai, Kevin, et al.
Published: (2025)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
by: Ichihara, Yuki, et al.
Published: (2025)
by: Ichihara, Yuki, et al.
Published: (2025)
The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
by: Sahoo, Subramanyam, et al.
Published: (2025)
by: Sahoo, Subramanyam, et al.
Published: (2025)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
by: Beigi, Mohammad, et al.
Published: (2026)
by: Beigi, Mohammad, et al.
Published: (2026)
Inference-Time Reward Hacking in Large Language Models
by: Khalaf, Hadi, et al.
Published: (2025)
by: Khalaf, Hadi, et al.
Published: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
by: Wang, Ye, et al.
Published: (2026)
by: Wang, Ye, et al.
Published: (2026)
Similar Items
-
Reward Model Generalization for Compute-Aware Test-Time Reasoning
by: Song, Zeen, et al.
Published: (2025) -
Hacking Task Confounder in Meta-Learning
by: Wang, Jingyao, et al.
Published: (2023) -
Reward Shaping to Mitigate Reward Hacking in RLHF
by: Fu, Jiayi, et al.
Published: (2025) -
Reward Hacking Mitigation using Verifiable Composite Rewards
by: Tarek, Mirza Farhan Bin, et al.
Published: (2025) -
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
by: Hatgis-Kessell, Stephane, et al.
Published: (2025)