InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Miao, Yuchun, Zhang, Sen, Ding, Liang, Bao, Rong, Zhang, Lefei, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
von: Miao, Yuchun, et al.
Veröffentlicht: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
Information-Theoretic Reward Decomposition for Generalizable RLHF
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
von: Mao, Liyuan, et al.
Veröffentlicht: (2025)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
von: Zhang, Chuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Chuheng, et al.
Veröffentlicht: (2024)
Reward Hacking Mitigation using Verifiable Composite Rewards
von: Tarek, Mirza Farhan Bin, et al.
Veröffentlicht: (2025)
von: Tarek, Mirza Farhan Bin, et al.
Veröffentlicht: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
von: Hatgis-Kessell, Stephane, et al.
Veröffentlicht: (2025)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
Mitigating Reward Over-Optimization in RLHF via Behavior-Supported Regularization
von: Dai, Juntao, et al.
Veröffentlicht: (2025)
von: Dai, Juntao, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
von: Singha, Disha
Veröffentlicht: (2026)
von: Singha, Disha
Veröffentlicht: (2026)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
von: Wang, Chaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Chaoqi, et al.
Veröffentlicht: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
A Theoretical Framework for Partially Observed Reward-States in RLHF
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
von: Kausik, Chinmaya, et al.
Veröffentlicht: (2024)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2026)
RM-R1: Reward Modeling as Reasoning
von: Chen, Xiusi, et al.
Veröffentlicht: (2025)
von: Chen, Xiusi, et al.
Veröffentlicht: (2025)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2024)
von: Laidlaw, Cassidy, et al.
Veröffentlicht: (2024)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
Intra-Trajectory Consistency for Reward Modeling
von: Zhou, Chaoyang, et al.
Veröffentlicht: (2025)
von: Zhou, Chaoyang, et al.
Veröffentlicht: (2025)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
von: Roth, Amit, et al.
Veröffentlicht: (2026)
von: Roth, Amit, et al.
Veröffentlicht: (2026)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
von: Liang, Yu, et al.
Veröffentlicht: (2026)
von: Liang, Yu, et al.
Veröffentlicht: (2026)
When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR
von: Miao, Yuchun, et al.
Veröffentlicht: (2026)
von: Miao, Yuchun, et al.
Veröffentlicht: (2026)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
von: Duan, Kaiwen, et al.
Veröffentlicht: (2025)
von: Duan, Kaiwen, et al.
Veröffentlicht: (2025)
AgentRM: Enhancing Agent Generalization with Reward Modeling
von: Xia, Yu, et al.
Veröffentlicht: (2025)
von: Xia, Yu, et al.
Veröffentlicht: (2025)
ALaRM: Align Language Models via Hierarchical Rewards Modeling
von: Lai, Yuhang, et al.
Veröffentlicht: (2024)
von: Lai, Yuhang, et al.
Veröffentlicht: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
GARDO: Reinforcing Diffusion Models without Reward Hacking
von: He, Haoran, et al.
Veröffentlicht: (2025)
von: He, Haoran, et al.
Veröffentlicht: (2025)
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
von: Farquhar, Sebastian, et al.
Veröffentlicht: (2025)
von: Farquhar, Sebastian, et al.
Veröffentlicht: (2025)
Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
von: Ouyang, Sheng, et al.
Veröffentlicht: (2025)
von: Ouyang, Sheng, et al.
Veröffentlicht: (2025)
UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
von: Fu, Lingling, et al.
Veröffentlicht: (2025)
von: Fu, Lingling, et al.
Veröffentlicht: (2025)
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
von: Jian, Ai, et al.
Veröffentlicht: (2025)
von: Jian, Ai, et al.
Veröffentlicht: (2025)
ProgRM: Build Better GUI Agents with Progress Rewards
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
von: Zhang, Danyang, et al.
Veröffentlicht: (2025)
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
von: Rashidinejad, Paria, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
von: Miao, Yuchun, et al.
Veröffentlicht: (2025) -
The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
von: Miao, Yuchun, et al.
Veröffentlicht: (2025) -
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025) -
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024) -
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)