Explaining and Preventing Alignment Collapse in Iterative RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Gauthier, Etienne, Bach, Francis, Jordan, Michael I. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Backward Conformal Prediction
by: Gauthier, Etienne, et al.
Published: (2025)
by: Gauthier, Etienne, et al.
Published: (2025)
E-Values Expand the Scope of Conformal Prediction
by: Gauthier, Etienne, et al.
Published: (2025)
by: Gauthier, Etienne, et al.
Published: (2025)
Statistical Collusion by Collectives on Learning Platforms
by: Gauthier, Etienne, et al.
Published: (2025)
by: Gauthier, Etienne, et al.
Published: (2025)
Adaptive Coverage Policies in Conformal Prediction
by: Gauthier, Etienne, et al.
Published: (2025)
by: Gauthier, Etienne, et al.
Published: (2025)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
by: Zhu, Banghua, et al.
Published: (2024)
by: Zhu, Banghua, et al.
Published: (2024)
Optimal Design for Reward Modeling in RLHF
by: Scheid, Antoine, et al.
Published: (2024)
by: Scheid, Antoine, et al.
Published: (2024)
Anytime Detection of Strategic Deviations in Multi-Agent Systems
by: Gauthier, Etienne, et al.
Published: (2026)
by: Gauthier, Etienne, et al.
Published: (2026)
Super-Level-Set Regression: Conditional Quantiles via Volume Minimization
by: Braun, Sacha, et al.
Published: (2026)
by: Braun, Sacha, et al.
Published: (2026)
Mitigating the Alignment Tax of RLHF
by: Lin, Yong, et al.
Published: (2023)
by: Lin, Yong, et al.
Published: (2023)
CalArena: A Large-Scale Post-Hoc Calibration Benchmark
by: Berta, Eugène, et al.
Published: (2026)
by: Berta, Eugène, et al.
Published: (2026)
Rethinking Early Stopping: Refine, Then Calibrate
by: Berta, Eugène, et al.
Published: (2025)
by: Berta, Eugène, et al.
Published: (2025)
Conditional Coverage Diagnostics for Conformal Prediction
by: Braun, Sacha, et al.
Published: (2025)
by: Braun, Sacha, et al.
Published: (2025)
Structured Matrix Scaling for Multi-Class Calibration
by: Berta, Eugène, et al.
Published: (2025)
by: Berta, Eugène, et al.
Published: (2025)
A Variational Estimator for $L_p$ Calibration Errors
by: Berta, Eugène, et al.
Published: (2026)
by: Berta, Eugène, et al.
Published: (2026)
Anchored Alignment: Preventing Positional Collapse in Multimodal Recommender Systems
by: Jeong, Yonghun, et al.
Published: (2026)
by: Jeong, Yonghun, et al.
Published: (2026)
Reward Model Overoptimisation in Iterated RLHF
by: Wolf, Lorenz, et al.
Published: (2025)
by: Wolf, Lorenz, et al.
Published: (2025)
Multivariate Standardized Residuals for Conformal Prediction
by: Braun, Sacha, et al.
Published: (2025)
by: Braun, Sacha, et al.
Published: (2025)
Minimum Volume Conformal Sets for Multivariate Regression
by: Braun, Sacha, et al.
Published: (2025)
by: Braun, Sacha, et al.
Published: (2025)
Beyond RLHF: A Unified Theoretical Framework of Alignment
by: Yun, Jihun, et al.
Published: (2025)
by: Yun, Jihun, et al.
Published: (2025)
How Sampling Shapes LLM Alignment: From One-Shot Optima to Iterative Dynamics
by: Chen, Yurong, et al.
Published: (2026)
by: Chen, Yurong, et al.
Published: (2026)
Solving the Inverse Alignment Problem for Efficient RLHF
by: Krishna, Shambhavi, et al.
Published: (2024)
by: Krishna, Shambhavi, et al.
Published: (2024)
Towards Reliable Alignment: Uncertainty-aware RLHF
by: Banerjee, Debangshu, et al.
Published: (2024)
by: Banerjee, Debangshu, et al.
Published: (2024)
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
by: Chegini, Atoosa, et al.
Published: (2024)
by: Chegini, Atoosa, et al.
Published: (2024)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
by: Sahoo, Subramanyam, et al.
Published: (2025)
by: Sahoo, Subramanyam, et al.
Published: (2025)
On the Algorithmic Bias of Aligning Large Language Models with RLHF: Preference Collapse and Matching Regularization
by: Xiao, Jiancong, et al.
Published: (2024)
by: Xiao, Jiancong, et al.
Published: (2024)
On the Effectiveness of the z-Transform Method in Quadratic Optimization
by: Bach, Francis
Published: (2025)
by: Bach, Francis
Published: (2025)
A Convex Loss Function for Set Prediction with Optimal Trade-offs Between Size and Conditional Coverage
by: Bach, Francis
Published: (2025)
by: Bach, Francis
Published: (2025)
Why Is RLHF Alignment Shallow? A Gradient Analysis
by: Young, Robin
Published: (2026)
by: Young, Robin
Published: (2026)
Preventing Model Collapse in Gaussian Process Latent Variable Models
by: Li, Ying, et al.
Published: (2024)
by: Li, Ying, et al.
Published: (2024)
Mutual Information Collapse Explains Disentanglement Failure in $β$-VAEs
by: Vu, Minh, et al.
Published: (2026)
by: Vu, Minh, et al.
Published: (2026)
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
by: Sakamoto, Keitaro, et al.
Published: (2025)
by: Sakamoto, Keitaro, et al.
Published: (2025)
Golden Ratio Weighting Prevents Model Collapse
by: He, Hengzhi, et al.
Published: (2025)
by: He, Hengzhi, et al.
Published: (2025)
RLHF Fine-Tuning of LLMs for Alignment with Implicit User Feedback in Conversational Recommenders
by: Yang, Zhongheng, et al.
Published: (2025)
by: Yang, Zhongheng, et al.
Published: (2025)
A Spectral Framework for Closed-Form Relative Density Estimation
by: Bach, Francis
Published: (2026)
by: Bach, Francis
Published: (2026)
Enhanced Feature Learning via Regularisation: Integrating Neural Networks and Kernel Methods
by: Follain, Bertille, et al.
Published: (2024)
by: Follain, Bertille, et al.
Published: (2024)
Sampling Binary Data by Denoising through Score Functions
by: Bach, Francis, et al.
Published: (2025)
by: Bach, Francis, et al.
Published: (2025)
Optimal Denoising in Score-Based Generative Models: The Role of Data Regularity
by: Beyler, Eliot, et al.
Published: (2025)
by: Beyler, Eliot, et al.
Published: (2025)
Convergence of Deterministic and Stochastic Diffusion-Model Samplers: A Simple Analysis in Wasserstein Distance
by: Beyler, Eliot, et al.
Published: (2025)
by: Beyler, Eliot, et al.
Published: (2025)
When Models Don't Collapse: On the Consistency of Iterative MLE
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
APPA: Adaptive Preference Pluralistic Alignment for Fair Federated RLHF of LLMs
by: Srewa, Mahmoud, et al.
Published: (2026)
by: Srewa, Mahmoud, et al.
Published: (2026)
Similar Items
-
Backward Conformal Prediction
by: Gauthier, Etienne, et al.
Published: (2025) -
E-Values Expand the Scope of Conformal Prediction
by: Gauthier, Etienne, et al.
Published: (2025) -
Statistical Collusion by Collectives on Learning Platforms
by: Gauthier, Etienne, et al.
Published: (2025) -
Adaptive Coverage Policies in Conformal Prediction
by: Gauthier, Etienne, et al.
Published: (2025) -
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
by: Zhu, Banghua, et al.
Published: (2024)