On Goodhart's law, with an application to value alignment
Fuente:
arXiv
Saved in:
| Main Authors: | El-Mhamdi, El-Mahdi, Hoang, Lê-Nguyên |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Strong, Weak and Benign Goodhart's law. An independence-free and paradigm-agnostic formalisation
by: Majka, Adrien, et al.
Published: (2025)
by: Majka, Adrien, et al.
Published: (2025)
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
by: Bareilles, Gilles, et al.
Published: (2026)
by: Bareilles, Gilles, et al.
Published: (2026)
On Monotonicity in AI Alignment
by: Bareilles, Gilles, et al.
Published: (2025)
by: Bareilles, Gilles, et al.
Published: (2025)
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
by: Firdoussi, Aymane El, et al.
Published: (2024)
by: Firdoussi, Aymane El, et al.
Published: (2024)
A Case for Specialisation in Non-Human Entities
by: El-Mhamdi, El-Mahdi, et al.
Published: (2025)
by: El-Mhamdi, El-Mahdi, et al.
Published: (2025)
Convergence of Statistical Estimators via Mutual Information Bounds
by: Khribch, El Mahdi, et al.
Published: (2024)
by: Khribch, El Mahdi, et al.
Published: (2024)
Class conditional conformal prediction for multiple inputs by p-value aggregation
by: Fermanian, Jean-Baptiste, et al.
Published: (2025)
by: Fermanian, Jean-Baptiste, et al.
Published: (2025)
Revisiting Incremental Stochastic Majorization-Minimization Algorithms with Applications to Mixture of Experts
by: Tran, TrungKhang, et al.
Published: (2026)
by: Tran, TrungKhang, et al.
Published: (2026)
Decoupled Continuous-Time Reinforcement Learning via Hamiltonian Flow
by: Nguyen, Minh
Published: (2026)
by: Nguyen, Minh
Published: (2026)
A Differential and Pointwise Control Approach to Reinforcement Learning
by: Nguyen, Minh, et al.
Published: (2024)
by: Nguyen, Minh, et al.
Published: (2024)
Influence functions and regularity tangents for efficient active learning
by: Eaton, Frederik
Published: (2024)
by: Eaton, Frederik
Published: (2024)
From Spikes to Heavy Tails: Unveiling the Spectral Evolution of Neural Networks
by: Kothapalli, Vignesh, et al.
Published: (2024)
by: Kothapalli, Vignesh, et al.
Published: (2024)
Generalising realisability in statistical learning theory under epistemic uncertainty
by: Cuzzolin, Fabio
Published: (2024)
by: Cuzzolin, Fabio
Published: (2024)
Training Implicit Generative Models via an Invariant Statistical Loss
by: de Frutos, José Manuel, et al.
Published: (2024)
by: de Frutos, José Manuel, et al.
Published: (2024)
U-Nets as Belief Propagation: Efficient Classification, Denoising, and Diffusion in Generative Hierarchical Models
by: Mei, Song
Published: (2024)
by: Mei, Song
Published: (2024)
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
by: Rajendran, Goutham, et al.
Published: (2024)
by: Rajendran, Goutham, et al.
Published: (2024)
Adaptive Sample Aggregation In Transfer Learning
by: Hanneke, Steve, et al.
Published: (2024)
by: Hanneke, Steve, et al.
Published: (2024)
A Statistical Analysis of Deep Federated Learning for Intrinsically Low-dimensional Data
by: Chakraborty, Saptarshi, et al.
Published: (2024)
by: Chakraborty, Saptarshi, et al.
Published: (2024)
Diffusion Posterior Sampling is Computationally Intractable
by: Gupta, Shivam, et al.
Published: (2024)
by: Gupta, Shivam, et al.
Published: (2024)
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
by: Javanmard, Adel, et al.
Published: (2024)
by: Javanmard, Adel, et al.
Published: (2024)
O(d/T) Convergence Theory for Diffusion Probabilistic Models under Minimal Assumptions
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Enhancing Conformal Prediction Using E-Test Statistics
by: Balinsky, A. A., et al.
Published: (2024)
by: Balinsky, A. A., et al.
Published: (2024)
Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks
by: Fu, Hengyu, et al.
Published: (2024)
by: Fu, Hengyu, et al.
Published: (2024)
Adapting to Unknown Low-Dimensional Structures in Score-Based Diffusion Models
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Towards Bayesian Data Selection
by: Rodemann, Julian
Published: (2024)
by: Rodemann, Julian
Published: (2024)
Counterfactual Generative Modeling with Variational Causal Inference
by: Wu, Yulun, et al.
Published: (2024)
by: Wu, Yulun, et al.
Published: (2024)
Scaling Laws in Linear Regression: Compute, Parameters, and Data
by: Lin, Licong, et al.
Published: (2024)
by: Lin, Licong, et al.
Published: (2024)
Neural Networks Generalize on Low Complexity Data
by: Chatterjee, Sourav, et al.
Published: (2024)
by: Chatterjee, Sourav, et al.
Published: (2024)
On the Statistical Properties of Generative Adversarial Models for Low Intrinsic Data Dimension
by: Chakraborty, Saptarshi, et al.
Published: (2024)
by: Chakraborty, Saptarshi, et al.
Published: (2024)
Online Learning with Unknown Constraints
by: Sridharan, Karthik, et al.
Published: (2024)
by: Sridharan, Karthik, et al.
Published: (2024)
Is Behavior Cloning All You Need? Understanding Horizon in Imitation Learning
by: Foster, Dylan J., et al.
Published: (2024)
by: Foster, Dylan J., et al.
Published: (2024)
A comparative study of conformal prediction methods for valid uncertainty quantification in machine learning
by: Dewolf, Nicolas
Published: (2024)
by: Dewolf, Nicolas
Published: (2024)
Optimal rates for density and mode estimation with expand-and-sparsify representations
by: Sinha, Kaushik, et al.
Published: (2026)
by: Sinha, Kaushik, et al.
Published: (2026)
Efficient Knowledge Distillation via Curriculum Extraction
by: Gupta, Shivam, et al.
Published: (2025)
by: Gupta, Shivam, et al.
Published: (2025)
Risk Analysis and Design Against Adversarial Actions
by: Campi, Marco C., et al.
Published: (2025)
by: Campi, Marco C., et al.
Published: (2025)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
by: Su, Xun, et al.
Published: (2025)
by: Su, Xun, et al.
Published: (2025)
A Theory of the Mechanics of Information: Generalization Through Measurement of Uncertainty (Learning is Measuring)
by: Hazard, Christopher J., et al.
Published: (2025)
by: Hazard, Christopher J., et al.
Published: (2025)
Solving a Research Problem in Mathematical Statistics with AI Assistance
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
Identifiability of Potentially Degenerate Gaussian Mixture Models With Piecewise Affine Mixing
by: Xu, Danru, et al.
Published: (2026)
by: Xu, Danru, et al.
Published: (2026)
Cross-regularization: Adaptive Model Complexity through Validation Gradients
by: Brito, Carlos Stein
Published: (2025)
by: Brito, Carlos Stein
Published: (2025)
Similar Items
-
The Strong, Weak and Benign Goodhart's law. An independence-free and paradigm-agnostic formalisation
by: Majka, Adrien, et al.
Published: (2025) -
Byzantine Machine Learning: MultiKrum and an optimal notion of robustness
by: Bareilles, Gilles, et al.
Published: (2026) -
On Monotonicity in AI Alignment
by: Bareilles, Gilles, et al.
Published: (2025) -
Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory
by: Firdoussi, Aymane El, et al.
Published: (2024) -
A Case for Specialisation in Non-Human Entities
by: El-Mhamdi, El-Mahdi, et al.
Published: (2025)