Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Qingyue, Ji, Kaixuan, Zhao, Heyang, Zhang, Tong, Gu, Quanquan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026)
by: Zhao, Qingyue, et al.
Published: (2026)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
by: Zhao, Heyang, et al.
Published: (2024)
by: Zhao, Heyang, et al.
Published: (2024)
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
Statistical Inference for Misspecified Contextual Bandits
by: Guo, Yongyi, et al.
Published: (2025)
by: Guo, Yongyi, et al.
Published: (2025)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025)
by: He, Jiafan, et al.
Published: (2025)
Feel-Good Thompson Sampling for Contextual Dueling Bandits
by: Li, Xuheng, et al.
Published: (2024)
by: Li, Xuheng, et al.
Published: (2024)
Distribution-consistency Structural Causal Models
by: Gong, Heyang, et al.
Published: (2024)
by: Gong, Heyang, et al.
Published: (2024)
Conformal Policy Control
by: Prinster, Drew, et al.
Published: (2026)
by: Prinster, Drew, et al.
Published: (2026)
Reasoning with Sampling: Cutting at Decision Points
by: Zhou, Felix, et al.
Published: (2026)
by: Zhou, Felix, et al.
Published: (2026)
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
by: Chen, Zhirui, et al.
Published: (2024)
by: Chen, Zhirui, et al.
Published: (2024)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
by: Chen, Zixiang, et al.
Published: (2025)
by: Chen, Zixiang, et al.
Published: (2025)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
by: Boudart, Pierre, et al.
Published: (2025)
by: Boudart, Pierre, et al.
Published: (2025)
LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning
by: Cao, Junyu, et al.
Published: (2026)
by: Cao, Junyu, et al.
Published: (2026)
Navigating Sparsities in High-Dimensional Linear Contextual Bandits
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes
by: Lu, Miao, et al.
Published: (2022)
by: Lu, Miao, et al.
Published: (2022)
Smooth Non-Stationary Bandits
by: Jia, Su, et al.
Published: (2023)
by: Jia, Su, et al.
Published: (2023)
Reinforcement Learning from Human Feedback with Active Queries
by: Ji, Kaixuan, et al.
Published: (2024)
by: Ji, Kaixuan, et al.
Published: (2024)
Robust Offline Reinforcement Learning with Linearly Structured f-Divergence Regularization
by: Tang, Cheng, et al.
Published: (2024)
by: Tang, Cheng, et al.
Published: (2024)
Ordinary Least Squares is a Special Case of Transformer
by: Tan, Xiaojun, et al.
Published: (2026)
by: Tan, Xiaojun, et al.
Published: (2026)
Transfer Learning for Contextual Multi-armed Bandits
by: Cai, Changxiao, et al.
Published: (2022)
by: Cai, Changxiao, et al.
Published: (2022)
Optimizing Warfarin Dosing Using Contextual Bandit: An Offline Policy Learning and Evaluation Method
by: Huang, Yong, et al.
Published: (2024)
by: Huang, Yong, et al.
Published: (2024)
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
by: Fan, Zhiyuan, et al.
Published: (2025)
by: Fan, Zhiyuan, et al.
Published: (2025)
Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon
by: Yu, Hao
Published: (2025)
by: Yu, Hao
Published: (2025)
A Statistical Analysis of Deep Federated Learning for Intrinsically Low-dimensional Data
by: Chakraborty, Saptarshi, et al.
Published: (2024)
by: Chakraborty, Saptarshi, et al.
Published: (2024)
Batched Nonparametric Contextual Bandits
by: Jiang, Rong, et al.
Published: (2024)
by: Jiang, Rong, et al.
Published: (2024)
Towards Bayesian Data Selection
by: Rodemann, Julian
Published: (2024)
by: Rodemann, Julian
Published: (2024)
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
by: Wu, Di, et al.
Published: (2026)
by: Wu, Di, et al.
Published: (2026)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
The Geometry of Benchmarks: A New Path Toward AGI
by: Chojecki, Przemyslaw
Published: (2025)
by: Chojecki, Przemyslaw
Published: (2025)
Characteristic Learning for Provable One Step Generation
by: Ding, Zhao, et al.
Published: (2024)
by: Ding, Zhao, et al.
Published: (2024)
Labels or Preferences? Budget-Constrained Learning with Human Judgments over AI-Generated Outputs
by: Dong, Zihan, et al.
Published: (2026)
by: Dong, Zihan, et al.
Published: (2026)
Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Risk Analysis and Design Against Adversarial Actions
by: Campi, Marco C., et al.
Published: (2025)
by: Campi, Marco C., et al.
Published: (2025)
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
by: Rajendran, Goutham, et al.
Published: (2024)
by: Rajendran, Goutham, et al.
Published: (2024)
Compression, Generalization and Learning
by: Campi, Marco C., et al.
Published: (2023)
by: Campi, Marco C., et al.
Published: (2023)
Online Learning with Unknown Constraints
by: Sridharan, Karthik, et al.
Published: (2024)
by: Sridharan, Karthik, et al.
Published: (2024)
Adaptive Sample Aggregation In Transfer Learning
by: Hanneke, Steve, et al.
Published: (2024)
by: Hanneke, Steve, et al.
Published: (2024)
Similar Items
-
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026) -
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026) -
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026) -
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
by: Zhao, Heyang, et al.
Published: (2024) -
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)