Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Qingyue, Ji, Kaixuan, Zhao, Heyang, Gu, Quanquan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025)
by: Zhao, Qingyue, et al.
Published: (2025)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
by: Chen, Zhirui, et al.
Published: (2024)
by: Chen, Zhirui, et al.
Published: (2024)
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
by: Praharaj, Samya, et al.
Published: (2025)
by: Praharaj, Samya, et al.
Published: (2025)
Sharp Analysis for KL-Regularized Contextual Bandits and RLHF
by: Zhao, Heyang, et al.
Published: (2024)
by: Zhao, Heyang, et al.
Published: (2024)
Statistical Inference for Misspecified Contextual Bandits
by: Guo, Yongyi, et al.
Published: (2025)
by: Guo, Yongyi, et al.
Published: (2025)
A Diffusion Analysis of Policy Gradient for Stochastic Bandits
by: Lattimore, Tor
Published: (2026)
by: Lattimore, Tor
Published: (2026)
Concentrated Differential Privacy for Bandits
by: Azize, Achraf, et al.
Published: (2023)
by: Azize, Achraf, et al.
Published: (2023)
Universal time-series forecasting with mixture predictors
by: Ryabko, Daniil
Published: (2020)
by: Ryabko, Daniil
Published: (2020)
MESSY Estimation: Maximum-Entropy based Stochastic and Symbolic densitY Estimation
by: Tohme, Tony, et al.
Published: (2023)
by: Tohme, Tony, et al.
Published: (2023)
Generalization Bounds: Perspectives from Information Theory and PAC-Bayes
by: Hellström, Fredrik, et al.
Published: (2023)
by: Hellström, Fredrik, et al.
Published: (2023)
Decision Making in Changing Environments: Robustness, Query-Based Learning, and Differential Privacy
by: Chen, Fan, et al.
Published: (2025)
by: Chen, Fan, et al.
Published: (2025)
Prior-dependent analysis of posterior sampling reinforcement learning with function approximation
by: Li, Yingru, et al.
Published: (2024)
by: Li, Yingru, et al.
Published: (2024)
Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit
by: Zhang, Bohan, et al.
Published: (2025)
by: Zhang, Bohan, et al.
Published: (2025)
Fixed-Budget Differentially Private Best Arm Identification
by: Chen, Zhirui, et al.
Published: (2024)
by: Chen, Zhirui, et al.
Published: (2024)
Conformal Policy Control
by: Prinster, Drew, et al.
Published: (2026)
by: Prinster, Drew, et al.
Published: (2026)
Optimal Regret of Bernoulli Bandits under Global Differential Privacy
by: Azize, Achraf, et al.
Published: (2025)
by: Azize, Achraf, et al.
Published: (2025)
Online Clustering of Data Sequences with Bandit Information
by: Chandran, G Dhinesh, et al.
Published: (2025)
by: Chandran, G Dhinesh, et al.
Published: (2025)
Distribution-consistency Structural Causal Models
by: Gong, Heyang, et al.
Published: (2024)
by: Gong, Heyang, et al.
Published: (2024)
On Instability of Minimax Optimal Optimism-Based Bandit Algorithms
by: Praharaj, Samya, et al.
Published: (2025)
by: Praharaj, Samya, et al.
Published: (2025)
Navigating Sparsities in High-Dimensional Linear Contextual Bandits
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
The Geometry of Knowing: From Possibilistic Ignorance to Probabilistic Certainty -- A Measure-Theoretic Framework for Epistemic Convergence
by: Jah, Moriba Kemessia
Published: (2026)
by: Jah, Moriba Kemessia
Published: (2026)
Transport f divergences
by: Li, Wuchen
Published: (2025)
by: Li, Wuchen
Published: (2025)
Random Multiplexing
by: Liu, Lei, et al.
Published: (2025)
by: Liu, Lei, et al.
Published: (2025)
Enjoying Non-linearity in Multinomial Logistic Bandits: A Minimax-Optimal Algorithm
by: Boudart, Pierre, et al.
Published: (2025)
by: Boudart, Pierre, et al.
Published: (2025)
LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning
by: Cao, Junyu, et al.
Published: (2026)
by: Cao, Junyu, et al.
Published: (2026)
Reasoning with Sampling: Cutting at Decision Points
by: Zhou, Felix, et al.
Published: (2026)
by: Zhou, Felix, et al.
Published: (2026)
Statistical Complexity and Optimal Algorithms for Non-linear Ridge Bandits
by: Rajaraman, Nived, et al.
Published: (2023)
by: Rajaraman, Nived, et al.
Published: (2023)
Analyzing Shapley Additive Explanations to Understand Anomaly Detection Algorithm Behaviors and Their Complementarity
by: Levy, Jordan, et al.
Published: (2026)
by: Levy, Jordan, et al.
Published: (2026)
Entropy, concentration, and learning: a statistical mechanics primer
by: Balsubramani, Akshay
Published: (2024)
by: Balsubramani, Akshay
Published: (2024)
On the Separability of Information in Diffusion Models
by: Premkumar, Akhil
Published: (2025)
by: Premkumar, Akhil
Published: (2025)
Smooth Non-Stationary Bandits
by: Jia, Su, et al.
Published: (2023)
by: Jia, Su, et al.
Published: (2023)
Assouad, Fano, and Le Cam with Interaction: A Unifying Lower Bound Framework and Characterization for Bandit Learnability
by: Chen, Fan, et al.
Published: (2024)
by: Chen, Fan, et al.
Published: (2024)
Guaranteed Recovery of Unambiguous Clusters
by: Mazooji, Kayvon, et al.
Published: (2025)
by: Mazooji, Kayvon, et al.
Published: (2025)
Empirical Risk Minimization with Relative Entropy Regularization
by: Perlaza, Samir M., et al.
Published: (2022)
by: Perlaza, Samir M., et al.
Published: (2022)
Total Variation Rates for Riemannian Flow Matching
by: Guan, Yunrui, et al.
Published: (2026)
by: Guan, Yunrui, et al.
Published: (2026)
Ordinary Least Squares is a Special Case of Transformer
by: Tan, Xiaojun, et al.
Published: (2026)
by: Tan, Xiaojun, et al.
Published: (2026)
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
by: Ryu, J. Jon, et al.
Published: (2025)
by: Ryu, J. Jon, et al.
Published: (2025)
Batched Nonparametric Contextual Bandits
by: Jiang, Rong, et al.
Published: (2024)
by: Jiang, Rong, et al.
Published: (2024)
Similar Items
-
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026) -
Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits
by: Zhao, Qingyue, et al.
Published: (2025) -
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026) -
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
by: Chen, Zhirui, et al.
Published: (2024) -
Avoiding the Price of Adaptivity: Inference in Linear Contextual Bandits via Stability
by: Praharaj, Samya, et al.
Published: (2025)