KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
Fuente:
arXiv
Saved in:
| Main Authors: | Aminian, Gholamali, Asadi, Amir R., Shenfeld, Idan, Mroueh, Youssef |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
Information Theoretic Guarantees For Policy Alignment In Large Language Models
by: Mroueh, Youssef
Published: (2024)
by: Mroueh, Youssef
Published: (2024)
Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
by: Mroueh, Youssef
Published: (2025)
by: Mroueh, Youssef
Published: (2025)
Generalization and Robustness of the Tilted Empirical Risk
by: Aminian, Gholamali, et al.
Published: (2024)
by: Aminian, Gholamali, et al.
Published: (2024)
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
by: Zhao, Heyang, et al.
Published: (2024)
by: Zhao, Heyang, et al.
Published: (2024)
Understanding Transfer Learning via Mean-field Analysis
by: Aminian, Gholamali, et al.
Published: (2024)
by: Aminian, Gholamali, et al.
Published: (2024)
UPL: Uncertainty-aware Pseudo-labeling for Imbalance Transductive Node Classification
by: Teimuri, Mohammad T., et al.
Published: (2025)
by: Teimuri, Mohammad T., et al.
Published: (2025)
Private Synthetic Graph Generation and Fused Gromov-Wasserstein Distance
by: Wirth, Leoni Carla, et al.
Published: (2025)
by: Wirth, Leoni Carla, et al.
Published: (2025)
Unifying Stable Optimization and Reference Regularization in RLHF
by: He, Li, et al.
Published: (2026)
by: He, Li, et al.
Published: (2026)
RL's Razor: Why Online Reinforcement Learning Forgets Less
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
Offline and Online KL-Regularized RLHF under Differential Privacy
by: Wu, Yulian, et al.
Published: (2025)
by: Wu, Yulian, et al.
Published: (2025)
Generalization Error of $f$-Divergence Stabilized Algorithms via Duality
by: Daunas, Francisco, et al.
Published: (2025)
by: Daunas, Francisco, et al.
Published: (2025)
Value Augmented Sampling for Language Model Alignment and Personalization
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
Language Model Personalization via Reward Factorization
by: Shenfeld, Idan, et al.
Published: (2025)
by: Shenfeld, Idan, et al.
Published: (2025)
Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
by: Liu, Kezhao, et al.
Published: (2025)
by: Liu, Kezhao, et al.
Published: (2025)
Learning Algorithm Generalization Error Bounds via Auxiliary Distributions
by: Aminian, Gholamali, et al.
Published: (2022)
by: Aminian, Gholamali, et al.
Published: (2022)
On the Generalization and Robustness in Conditional Value-at-Risk
by: Mulumudi, Dinesh Karthik, et al.
Published: (2026)
by: Mulumudi, Dinesh Karthik, et al.
Published: (2026)
Exact Recovery in the Data Block Model
by: Asadi, Amir R., et al.
Published: (2026)
by: Asadi, Amir R., et al.
Published: (2026)
$f$-FUM: Federated Unlearning via min--max and $f$-divergence
by: Karimian, Radmehr, et al.
Published: (2026)
by: Karimian, Radmehr, et al.
Published: (2026)
Self-Distillation Enables Continual Learning
by: Shenfeld, Idan, et al.
Published: (2026)
by: Shenfeld, Idan, et al.
Published: (2026)
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
ReDiF: Reinforced Distillation for Few Step Diffusion
by: Tighkhorshid, Amirhossein, et al.
Published: (2025)
by: Tighkhorshid, Amirhossein, et al.
Published: (2025)
Generalization Error of Graph Neural Networks in the Mean-field Regime
by: Aminian, Gholamali, et al.
Published: (2024)
by: Aminian, Gholamali, et al.
Published: (2024)
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
by: Shenfeld, Idan, et al.
Published: (2023)
by: Shenfeld, Idan, et al.
Published: (2023)
An Algorithm for Computing the Capacity of Symmetrized KL Information for Discrete Channels
by: Chen, Haobo, et al.
Published: (2024)
by: Chen, Haobo, et al.
Published: (2024)
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
by: Geuter, Jonathan, et al.
Published: (2025)
by: Geuter, Jonathan, et al.
Published: (2025)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Robust Semi-supervised Learning via $f$-Divergence and $α$-Rényi Divergence
by: Aminian, Gholamali, et al.
Published: (2024)
by: Aminian, Gholamali, et al.
Published: (2024)
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
by: Behnamnia, Armin, et al.
Published: (2025)
by: Behnamnia, Armin, et al.
Published: (2025)
CliffSearch: Structured Agentic Co-Evolution over Theory and Code for Scientific Algorithm Discovery
by: Mroueh, Youssef, et al.
Published: (2026)
by: Mroueh, Youssef, et al.
Published: (2026)
KL-regularization Itself is Differentially Private in Bandits and RLHF
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
Multivariate Stochastic Dominance via Optimal Transport and Applications to Models Benchmarking
by: Rioux, Gabriel, et al.
Published: (2024)
by: Rioux, Gabriel, et al.
Published: (2024)
From Imitation to Refinement -- Residual RL for Precise Assembly
by: Ankile, Lars, et al.
Published: (2024)
by: Ankile, Lars, et al.
Published: (2024)
Hierarchical Maximum Entropy via the Renormalization Group
by: Asadi, Amir R.
Published: (2025)
by: Asadi, Amir R.
Published: (2025)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
by: Tang, Kenton, et al.
Published: (2026)
by: Tang, Kenton, et al.
Published: (2026)
Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-End
by: Hanneke, Steve, et al.
Published: (2026)
by: Hanneke, Steve, et al.
Published: (2026)
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
by: Wu, Di, et al.
Published: (2026)
by: Wu, Di, et al.
Published: (2026)
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
by: Pipano, Idan, et al.
Published: (2026)
by: Pipano, Idan, et al.
Published: (2026)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
by: Puri, Isha, et al.
Published: (2026)
by: Puri, Isha, et al.
Published: (2026)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
by: Hilel, Almog, et al.
Published: (2025)
by: Hilel, Almog, et al.
Published: (2025)
Similar Items
-
Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis
by: Aminian, Gholamali, et al.
Published: (2025) -
Information Theoretic Guarantees For Policy Alignment In Large Language Models
by: Mroueh, Youssef
Published: (2024) -
Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
by: Mroueh, Youssef
Published: (2025) -
Generalization and Robustness of the Tilted Empirical Risk
by: Aminian, Gholamali, et al.
Published: (2024) -
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
by: Zhao, Heyang, et al.
Published: (2024)