Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
Fuente:
arXiv
Guardado en:
| Autores principales: | Kwa, Thomas, Thomas, Drake, Garriga-Alonso, Adrià |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
por: Gupta, Rohan, et al.
Publicado: (2024)
por: Gupta, Rohan, et al.
Publicado: (2024)
Investigating the Indirect Object Identification circuit in Mamba
por: Ensign, Danielle, et al.
Publicado: (2024)
por: Ensign, Danielle, et al.
Publicado: (2024)
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
por: Golechha, Satvik, et al.
Publicado: (2025)
por: Golechha, Satvik, et al.
Publicado: (2025)
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
por: Chanin, David, et al.
Publicado: (2026)
por: Chanin, David, et al.
Publicado: (2026)
KL-regularization Itself is Differentially Private in Bandits and RLHF
por: Zhang, Yizhou, et al.
Publicado: (2025)
por: Zhang, Yizhou, et al.
Publicado: (2025)
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
Adversarial Circuit Evaluation
por: de Bos, Niels uit, et al.
Publicado: (2024)
por: de Bos, Niels uit, et al.
Publicado: (2024)
Interpreting Emergent Planning in Model-Free Reinforcement Learning
por: Bush, Thomas, et al.
Publicado: (2025)
por: Bush, Thomas, et al.
Publicado: (2025)
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
por: Chanin, David, et al.
Publicado: (2025)
por: Chanin, David, et al.
Publicado: (2025)
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
por: Zhao, Heyang, et al.
Publicado: (2024)
por: Zhao, Heyang, et al.
Publicado: (2024)
Biases in the Blind Spot: Detecting What LLMs Fail to Mention
por: Arcuschin, Iván, et al.
Publicado: (2026)
por: Arcuschin, Iván, et al.
Publicado: (2026)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
por: Taufeeque, Mohammad, et al.
Publicado: (2025)
por: Taufeeque, Mohammad, et al.
Publicado: (2025)
On Goodhart's law, with an application to value alignment
por: El-Mhamdi, El-Mahdi, et al.
Publicado: (2024)
por: El-Mhamdi, El-Mahdi, et al.
Publicado: (2024)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
por: Tang, Kenton, et al.
Publicado: (2026)
por: Tang, Kenton, et al.
Publicado: (2026)
Offline and Online KL-Regularized RLHF under Differential Privacy
por: Wu, Yulian, et al.
Publicado: (2025)
por: Wu, Yulian, et al.
Publicado: (2025)
The PESQetarian: On the Relevance of Goodhart's Law for Speech Enhancement
por: de Oliveira, Danilo, et al.
Publicado: (2024)
por: de Oliveira, Danilo, et al.
Publicado: (2024)
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
por: Aminian, Gholamali, et al.
Publicado: (2025)
por: Aminian, Gholamali, et al.
Publicado: (2025)
Wasserstein KL-divergence for Gaussian distributions
por: Datar, Adwait, et al.
Publicado: (2025)
por: Datar, Adwait, et al.
Publicado: (2025)
Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
por: Liu, Kezhao, et al.
Publicado: (2025)
por: Liu, Kezhao, et al.
Publicado: (2025)
DiFR: Inference Verification Despite Nondeterminism
por: Karvonen, Adam, et al.
Publicado: (2025)
por: Karvonen, Adam, et al.
Publicado: (2025)
Learning Exceptional Subgroups by End-to-End Maximizing KL-divergence
por: Xu, Sascha, et al.
Publicado: (2024)
por: Xu, Sascha, et al.
Publicado: (2024)
On a few pitfalls in KL divergence gradient estimation for RL
por: Tang, Yunhao, et al.
Publicado: (2025)
por: Tang, Yunhao, et al.
Publicado: (2025)
A KL-regularization Framework for Learning to Plan with Adaptive Priors
por: Serra-Gomez, Álvaro, et al.
Publicado: (2025)
por: Serra-Gomez, Álvaro, et al.
Publicado: (2025)
Analyzing the Generalization and Reliability of Steering Vectors
por: Tan, Daniel, et al.
Publicado: (2024)
por: Tan, Daniel, et al.
Publicado: (2024)
EXACFS -- A CIL Method to mitigate Catastrophic Forgetting
por: Balasubramanian, S, et al.
Publicado: (2024)
por: Balasubramanian, S, et al.
Publicado: (2024)
The Strong, Weak and Benign Goodhart's law. An independence-free and paradigm-agnostic formalisation
por: Majka, Adrien, et al.
Publicado: (2025)
por: Majka, Adrien, et al.
Publicado: (2025)
Distributed, communication-efficient, and differentially private estimation of KL divergence
por: Scott, Mary, et al.
Publicado: (2024)
por: Scott, Mary, et al.
Publicado: (2024)
Are There Exceptions to Goodhart's Law? On the Moral Justification of Fairness-Aware Machine Learning
por: Weerts, Hilde, et al.
Publicado: (2022)
por: Weerts, Hilde, et al.
Publicado: (2022)
The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives
por: Moitra, Ankur, et al.
Publicado: (2026)
por: Moitra, Ankur, et al.
Publicado: (2026)
Improved AutoEncoder with LSTM module and KL divergence
por: Huang, Wei, et al.
Publicado: (2024)
por: Huang, Wei, et al.
Publicado: (2024)
Simulation-based Bayesian inference under model misspecification
por: Kelly, Ryan P., et al.
Publicado: (2025)
por: Kelly, Ryan P., et al.
Publicado: (2025)
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
por: Xiong, Wei, et al.
Publicado: (2023)
por: Xiong, Wei, et al.
Publicado: (2023)
Planning in a recurrent neural network that plays Sokoban
por: Taufeeque, Mohammad, et al.
Publicado: (2024)
por: Taufeeque, Mohammad, et al.
Publicado: (2024)
Do regularization methods for shortcut mitigation work as intended?
por: Hong, Haoyang, et al.
Publicado: (2025)
por: Hong, Haoyang, et al.
Publicado: (2025)
Gradient-based filtering under misspecification: Stability and error bounds
por: van Heel, Simon Donker, et al.
Publicado: (2025)
por: van Heel, Simon Donker, et al.
Publicado: (2025)
Rescuing double robustness: safe estimation under complete misspecification
por: Testa, Lorenzo, et al.
Publicado: (2025)
por: Testa, Lorenzo, et al.
Publicado: (2025)
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
por: Mukherjee, Arpan, et al.
Publicado: (2025)
por: Mukherjee, Arpan, et al.
Publicado: (2025)
Predicting and improving test-time scaling laws via reward tail-guided search
por: Li, Muheng, et al.
Publicado: (2026)
por: Li, Muheng, et al.
Publicado: (2026)
Variational f-divergence Minimization
por: Zhang, Mingtian, et al.
Publicado: (2019)
por: Zhang, Mingtian, et al.
Publicado: (2019)
Residual Policy Gradient: A Reward View of KL-regularized Objective
por: Wang, Pengcheng, et al.
Publicado: (2025)
por: Wang, Pengcheng, et al.
Publicado: (2025)
Ejemplares similares
-
InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques
por: Gupta, Rohan, et al.
Publicado: (2024) -
Investigating the Indirect Object Identification circuit in Mamba
por: Ensign, Danielle, et al.
Publicado: (2024) -
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
por: Golechha, Satvik, et al.
Publicado: (2025) -
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
por: Chanin, David, et al.
Publicado: (2026) -
KL-regularization Itself is Differentially Private in Bandits and RLHF
por: Zhang, Yizhou, et al.
Publicado: (2025)