The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sahoo, Subramanyam, Junkin, Jared |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Deepfake Detective: Interpreting Neural Forensics Through Sparse Features and Manifolds
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
von: Sahoo, Subramanyam
Veröffentlicht: (2025)
von: Sahoo, Subramanyam
Veröffentlicht: (2025)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
Causal Masking on Spatial Data: An Information-Theoretic Case for Learning Spatial Datasets with Unimodal Language Models
von: Junkin, Jared, et al.
Veröffentlicht: (2025)
von: Junkin, Jared, et al.
Veröffentlicht: (2025)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
The Double Life of Code World Models: Provably Unmasking Malicious Behavior Through Execution Traces
von: Sahoo, Subramanyam
Veröffentlicht: (2025)
von: Sahoo, Subramanyam
Veröffentlicht: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
von: Tarek, Mirza Farhan Bin, et al.
Veröffentlicht: (2025)
von: Tarek, Mirza Farhan Bin, et al.
Veröffentlicht: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
von: Liu, Zixuan, et al.
Veröffentlicht: (2026)
von: Liu, Zixuan, et al.
Veröffentlicht: (2026)
Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2023)
von: Eisenstein, Jacob, et al.
Veröffentlicht: (2023)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
von: Song, Ruike, et al.
Veröffentlicht: (2025)
von: Song, Ruike, et al.
Veröffentlicht: (2025)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
von: Yu, Zhuohao, et al.
Veröffentlicht: (2026)
von: Yu, Zhuohao, et al.
Veröffentlicht: (2026)
Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
von: Wu, Yusong, et al.
Veröffentlicht: (2025)
von: Wu, Yusong, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
von: Singha, Disha
Veröffentlicht: (2026)
von: Singha, Disha
Veröffentlicht: (2026)
UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
von: Fu, Lingling, et al.
Veröffentlicht: (2025)
von: Fu, Lingling, et al.
Veröffentlicht: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
von: Miao, Yuchun, et al.
Veröffentlicht: (2024)
von: Miao, Yuchun, et al.
Veröffentlicht: (2024)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
von: Wang, Songtao, et al.
Veröffentlicht: (2026)
von: Wang, Songtao, et al.
Veröffentlicht: (2026)
Boardwalk Empire: How Generative AI is Revolutionizing Economic Paradigms
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2024)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2024)
Defining and Characterizing Reward Hacking
von: Skalse, Joar, et al.
Veröffentlicht: (2022)
von: Skalse, Joar, et al.
Veröffentlicht: (2022)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
von: Wu, Rui, et al.
Veröffentlicht: (2026)
von: Wu, Rui, et al.
Veröffentlicht: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
von: Deng, Wenlong, et al.
Veröffentlicht: (2026)
von: Deng, Wenlong, et al.
Veröffentlicht: (2026)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2024)
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2024)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
von: Lian, Jiesong, et al.
Veröffentlicht: (2025)
von: Lian, Jiesong, et al.
Veröffentlicht: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
von: Roth, Amit, et al.
Veröffentlicht: (2026)
von: Roth, Amit, et al.
Veröffentlicht: (2026)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
von: Zhai, Kevin, et al.
Veröffentlicht: (2025)
von: Zhai, Kevin, et al.
Veröffentlicht: (2025)
MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
von: Farquhar, Sebastian, et al.
Veröffentlicht: (2025)
von: Farquhar, Sebastian, et al.
Veröffentlicht: (2025)
Inference-Time Reward Hacking in Large Language Models
von: Khalaf, Hadi, et al.
Veröffentlicht: (2025)
von: Khalaf, Hadi, et al.
Veröffentlicht: (2025)
Pedagogical Safety in Educational Reinforcement Learning: Formalizing and Detecting Reward Hacking in AI Tutoring Systems
von: Olukola, Oluseyi, et al.
Veröffentlicht: (2026)
von: Olukola, Oluseyi, et al.
Veröffentlicht: (2026)
Hacking Task Confounder in Meta-Learning
von: Wang, Jingyao, et al.
Veröffentlicht: (2023)
von: Wang, Jingyao, et al.
Veröffentlicht: (2023)
reward-lens: A Mechanistic Interpretability Library for Reward Models
von: Nadaf, Mohammed Suhail B
Veröffentlicht: (2026)
von: Nadaf, Mohammed Suhail B
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Deepfake Detective: Interpreting Neural Forensics Through Sparse Features and Manifolds
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025) -
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
von: Sahoo, Subramanyam
Veröffentlicht: (2026) -
The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
von: Sahoo, Subramanyam
Veröffentlicht: (2025) -
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026) -
Causal Masking on Spatial Data: An Information-Theoretic Case for Learning Spatial Datasets with Unimodal Language Models
von: Junkin, Jared, et al.
Veröffentlicht: (2025)