Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking
Fuente:
arXiv
Guardado en:
| Autores principales: | Eisenstein, Jacob, Nagpal, Chirag, Agarwal, Alekh, Beirami, Ahmad, D'Amour, Alex, Dvijotham, DJ, Fisch, Adam, Heller, Katherine, Pfohl, Stephen, Ramachandran, Deepak, Shaw, Peter, Berant, Jonathan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Robust Preference Optimization through Reward Model Distillation
por: Fisch, Adam, et al.
Publicado: (2024)
por: Fisch, Adam, et al.
Publicado: (2024)
Theoretical guarantees on the best-of-n alignment policy
por: Beirami, Ahmad, et al.
Publicado: (2024)
por: Beirami, Ahmad, et al.
Publicado: (2024)
Transforming and Combining Rewards for Aligning Large Language Models
por: Wang, Zihao, et al.
Publicado: (2024)
por: Wang, Zihao, et al.
Publicado: (2024)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
por: Setlur, Amrith, et al.
Publicado: (2024)
por: Setlur, Amrith, et al.
Publicado: (2024)
Mitigating Preference Hacking in Policy Optimization with Pessimism
por: Gupta, Dhawal, et al.
Publicado: (2025)
por: Gupta, Dhawal, et al.
Publicado: (2025)
Cost-Optimal Active AI Model Evaluation
por: Angelopoulos, Anastasios N., et al.
Publicado: (2025)
por: Angelopoulos, Anastasios N., et al.
Publicado: (2025)
Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
por: Lum, Kristian, et al.
Publicado: (2024)
por: Lum, Kristian, et al.
Publicado: (2024)
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment
por: Wu, Zhaofeng, et al.
Publicado: (2024)
por: Wu, Zhaofeng, et al.
Publicado: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Reward Model Ensembles Help Mitigate Overoptimization
por: Coste, Thomas, et al.
Publicado: (2023)
por: Coste, Thomas, et al.
Publicado: (2023)
Reward Hacking Mitigation using Verifiable Composite Rewards
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
por: Tarek, Mirza Farhan Bin, et al.
Publicado: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
por: Singha, Disha
Publicado: (2026)
por: Singha, Disha
Publicado: (2026)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
por: Beigi, Mohammad, et al.
Publicado: (2026)
por: Beigi, Mohammad, et al.
Publicado: (2026)
RRM: Robust Reward Model Training Mitigates Reward Hacking
por: Liu, Tianqi, et al.
Publicado: (2024)
por: Liu, Tianqi, et al.
Publicado: (2024)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
por: Ye, Chenlu, et al.
Publicado: (2025)
por: Ye, Chenlu, et al.
Publicado: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
por: Lian, Jiesong, et al.
Publicado: (2025)
por: Lian, Jiesong, et al.
Publicado: (2025)
Expected Reward Prediction, with Applications to Model Routing
por: Hasanaliyev, Kenan, et al.
Publicado: (2026)
por: Hasanaliyev, Kenan, et al.
Publicado: (2026)
Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
por: Pfohl, Stephen R., et al.
Publicado: (2025)
por: Pfohl, Stephen R., et al.
Publicado: (2025)
Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
por: Miao, Yuchun, et al.
Publicado: (2025)
por: Miao, Yuchun, et al.
Publicado: (2025)
Robust Optimization for Mitigating Reward Hacking with Correlated Proxies
por: Liu, Zixuan, et al.
Publicado: (2026)
por: Liu, Zixuan, et al.
Publicado: (2026)
MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games
por: Eisenstein, Jacob, et al.
Publicado: (2026)
por: Eisenstein, Jacob, et al.
Publicado: (2026)
UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
por: Fu, Lingling, et al.
Publicado: (2025)
por: Fu, Lingling, et al.
Publicado: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
por: Miao, Yuchun, et al.
Publicado: (2024)
por: Miao, Yuchun, et al.
Publicado: (2024)
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
por: Song, Ruike, et al.
Publicado: (2025)
por: Song, Ruike, et al.
Publicado: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
Utility-inspired Reward Transformations Improve Reinforcement Learning Training of Language Models
por: Maura-Rivero, Roberto-Rafael, et al.
Publicado: (2025)
por: Maura-Rivero, Roberto-Rafael, et al.
Publicado: (2025)
Proxy Methods for Domain Adaptation
por: Tsai, Katherine, et al.
Publicado: (2024)
por: Tsai, Katherine, et al.
Publicado: (2024)
InfAlign: Inference-aware language model alignment
por: Balashankar, Ananth, et al.
Publicado: (2024)
por: Balashankar, Ananth, et al.
Publicado: (2024)
Defining and Characterizing Reward Hacking
por: Skalse, Joar, et al.
Publicado: (2022)
por: Skalse, Joar, et al.
Publicado: (2022)
Copula-based Sensitivity Analysis for Multi-Treatment Causal Inference with Unobserved Confounding
por: Zheng, Jiajing, et al.
Publicado: (2021)
por: Zheng, Jiajing, et al.
Publicado: (2021)
Preference Models assume Proportional Hazards of Utilities
por: Nagpal, Chirag
Publicado: (2025)
por: Nagpal, Chirag
Publicado: (2025)
The Case for Globalizing Fairness: A Mixed Methods Study on Colonialism, AI, and Health in Africa
por: Asiedu, Mercy, et al.
Publicado: (2024)
por: Asiedu, Mercy, et al.
Publicado: (2024)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
por: Wu, Rui, et al.
Publicado: (2026)
por: Wu, Rui, et al.
Publicado: (2026)
From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism
por: Yu, Zhuohao, et al.
Publicado: (2026)
por: Yu, Zhuohao, et al.
Publicado: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
por: Deng, Wenlong, et al.
Publicado: (2026)
por: Deng, Wenlong, et al.
Publicado: (2026)
Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
por: Laidlaw, Cassidy, et al.
Publicado: (2024)
por: Laidlaw, Cassidy, et al.
Publicado: (2024)
Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation
por: Liu, Yifeng, et al.
Publicado: (2026)
por: Liu, Yifeng, et al.
Publicado: (2026)
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
por: Ferreira, Pedro, et al.
Publicado: (2025)
por: Ferreira, Pedro, et al.
Publicado: (2025)
Ejemplares similares
-
Robust Preference Optimization through Reward Model Distillation
por: Fisch, Adam, et al.
Publicado: (2024) -
Theoretical guarantees on the best-of-n alignment policy
por: Beirami, Ahmad, et al.
Publicado: (2024) -
Transforming and Combining Rewards for Aligning Large Language Models
por: Wang, Zihao, et al.
Publicado: (2024) -
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
por: Setlur, Amrith, et al.
Publicado: (2024) -
Mitigating Preference Hacking in Policy Optimization with Pessimism
por: Gupta, Dhawal, et al.
Publicado: (2025)