Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Heyang, Ye, Chenlu, Gu, Quanquan, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Logarithmic Regret for Online KL-Regularized Reinforcement Learning
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025)
by: Zhao, Qingyue, et al.
Published: (2025)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026)
by: Zhao, Qingyue, et al.
Published: (2026)
Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision Processes
by: Ye, Chenlu, et al.
Published: (2022)
by: Ye, Chenlu, et al.
Published: (2022)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Feel-Good Thompson Sampling for Contextual Dueling Bandits
by: Li, Xuheng, et al.
Published: (2024)
by: Li, Xuheng, et al.
Published: (2024)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
by: Ye, Chenlu, et al.
Published: (2025)
by: Ye, Chenlu, et al.
Published: (2025)
Towards Robust Model-Based Reinforcement Learning Against Adversarial Corruption
by: Ye, Chenlu, et al.
Published: (2024)
by: Ye, Chenlu, et al.
Published: (2024)
Corruption-Robust Offline Reinforcement Learning with General Function Approximation
by: Ye, Chenlu, et al.
Published: (2023)
by: Ye, Chenlu, et al.
Published: (2023)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025)
by: He, Jiafan, et al.
Published: (2025)
Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits
by: Li, Xuheng, et al.
Published: (2025)
by: Li, Xuheng, et al.
Published: (2025)
Nearly Optimal Algorithms for Contextual Dueling Bandits from Adversarial Feedback
by: Di, Qiwei, et al.
Published: (2024)
by: Di, Qiwei, et al.
Published: (2024)
KL-regularization Itself is Differentially Private in Bandits and RLHF
by: Zhang, Yizhou, et al.
Published: (2025)
by: Zhang, Yizhou, et al.
Published: (2025)
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
by: Xiong, Wei, et al.
Published: (2023)
by: Xiong, Wei, et al.
Published: (2023)
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
by: Zhao, Heyang, et al.
Published: (2023)
by: Zhao, Heyang, et al.
Published: (2023)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Offline and Online KL-Regularized RLHF under Differential Privacy
by: Wu, Yulian, et al.
Published: (2025)
by: Wu, Yulian, et al.
Published: (2025)
RIE-Greedy: Regularization-Induced Exploration for Contextual Bandits
by: Li, Tong, et al.
Published: (2026)
by: Li, Tong, et al.
Published: (2026)
Best-of-Majority: Minimax-Optimal Strategy for Pass@$k$ Inference Scaling
by: Di, Qiwei, et al.
Published: (2025)
by: Di, Qiwei, et al.
Published: (2025)
KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity
by: Aminian, Gholamali, et al.
Published: (2025)
by: Aminian, Gholamali, et al.
Published: (2025)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
by: Liu, Kezhao, et al.
Published: (2025)
by: Liu, Kezhao, et al.
Published: (2025)
Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence
by: Zhang, Shiyuan, et al.
Published: (2026)
by: Zhang, Shiyuan, et al.
Published: (2026)
A Sharp KL-Convergence Analysis for Diffusion Models under Minimal Assumptions
by: Jain, Nishant, et al.
Published: (2025)
by: Jain, Nishant, et al.
Published: (2025)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
Pure Exploration in Asynchronous Federated Bandits
by: Wang, Zichen, et al.
Published: (2023)
by: Wang, Zichen, et al.
Published: (2023)
Unifying Stable Optimization and Reference Regularization in RLHF
by: He, Li, et al.
Published: (2026)
by: He, Li, et al.
Published: (2026)
Contextual Bandits for Unbounded Context Distributions
by: Zhao, Puning, et al.
Published: (2024)
by: Zhao, Puning, et al.
Published: (2024)
Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization Term
by: Yue, Yun, et al.
Published: (2023)
by: Yue, Yun, et al.
Published: (2023)
Convergence of Score-Based Discrete Diffusion Models: A Discrete-Time Analysis
by: Zhang, Zikun, et al.
Published: (2024)
by: Zhang, Zikun, et al.
Published: (2024)
Differentially Private Kernelized Contextual Bandits
by: Pavlovic, Nikola, et al.
Published: (2025)
by: Pavlovic, Nikola, et al.
Published: (2025)
Online Bandit Learning with Offline Preference Data for Improved RLHF
by: Agnihotri, Akhil, et al.
Published: (2024)
by: Agnihotri, Akhil, et al.
Published: (2024)
ROCM: RLHF on consistency models
by: Shekhar, Shivanshu, et al.
Published: (2025)
by: Shekhar, Shivanshu, et al.
Published: (2025)
Generalisation of RLHF under Reward Shift and Clipped KL Regularisation
by: Tang, Kenton, et al.
Published: (2026)
by: Tang, Kenton, et al.
Published: (2026)
Contextual Linear Bandits with Delay as Payoff
by: Zhang, Mengxiao, et al.
Published: (2025)
by: Zhang, Mengxiao, et al.
Published: (2025)
Efficient Contextual Bandits with Uninformed Feedback Graphs
by: Zhang, Mengxiao, et al.
Published: (2024)
by: Zhang, Mengxiao, et al.
Published: (2024)
Sharpe Ratio-Guided Active Learning for Preference Optimization in RLHF
by: Belakaria, Syrine, et al.
Published: (2025)
by: Belakaria, Syrine, et al.
Published: (2025)
Sparse Nonparametric Contextual Bandits
by: Flynn, Hamish, et al.
Published: (2025)
by: Flynn, Hamish, et al.
Published: (2025)
Similar Items
-
Logarithmic Regret for Online KL-Regularized Reinforcement Learning
by: Zhao, Heyang, et al.
Published: (2025) -
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025) -
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026) -
Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision Processes
by: Ye, Chenlu, et al.
Published: (2022) -
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)