Kullback-Leibler Maillard Sampling for Multi-armed Bandits with Bounded Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Hao, Jun, Kwang-Sung, Zhang, Chicheng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
by: Qin, Hao, et al.
Published: (2025)
by: Qin, Hao, et al.
Published: (2025)
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
by: Jang, Kyoungseok, et al.
Published: (2024)
by: Jang, Kyoungseok, et al.
Published: (2024)
Taming the Monster Every Context: Complexity Measure and Unified Framework for Offline-Oracle Efficient Contextual Bandits
by: Qin, Hao, et al.
Published: (2026)
by: Qin, Hao, et al.
Published: (2026)
Transfer in Sequential Multi-armed Bandits via Reward Samples
by: R, Rahul N, et al.
Published: (2024)
by: R, Rahul N, et al.
Published: (2024)
Decoupled Kullback-Leibler Divergence Loss
by: Cui, Jiequan, et al.
Published: (2023)
by: Cui, Jiequan, et al.
Published: (2023)
Sharper Perturbed-Kullback-Leibler Exponential Tail Bounds for Beta and Dirichlet Distributions
by: Perrault, Pierre
Published: (2025)
by: Perrault, Pierre
Published: (2025)
Generalized Kullback-Leibler Divergence Loss
by: Cui, Jiequan, et al.
Published: (2025)
by: Cui, Jiequan, et al.
Published: (2025)
Physics-Informed Parametric Bandits for Beam Alignment in mmWave Communications
by: Qin, Hao, et al.
Published: (2025)
by: Qin, Hao, et al.
Published: (2025)
Kullback-Leibler Barycentre of Stochastic Processes
by: Jaimungal, Sebastian, et al.
Published: (2024)
by: Jaimungal, Sebastian, et al.
Published: (2024)
Connecting Jensen-Shannon and Kullback-Leibler Divergences: A New Bound for Representation Learning
by: Dorent, Reuben, et al.
Published: (2025)
by: Dorent, Reuben, et al.
Published: (2025)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
by: Lee, Junghyun, et al.
Published: (2023)
by: Lee, Junghyun, et al.
Published: (2023)
Beyond Task Diversity: Provable Representation Transfer for Sequential Multi-Task Linear Bandits
by: Duong, Thang, et al.
Published: (2025)
by: Duong, Thang, et al.
Published: (2025)
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
by: Lin, Wu, et al.
Published: (2025)
by: Lin, Wu, et al.
Published: (2025)
Risk Bounds for Mixture Density Estimation on Compact Domains via the $h$-Lifted Kullback--Leibler Divergence
by: Chong, Mark Chiu, et al.
Published: (2024)
by: Chong, Mark Chiu, et al.
Published: (2024)
Minimum Empirical Divergence for Sub-Gaussian Linear Bandits
by: Balagopalan, Kapilan, et al.
Published: (2024)
by: Balagopalan, Kapilan, et al.
Published: (2024)
Differentiable Annealed Importance Sampling Minimizes The Symmetrized Kullback-Leibler Divergence Between Initial and Target Distribution
by: Zenn, Johannes, et al.
Published: (2024)
by: Zenn, Johannes, et al.
Published: (2024)
Coarse-Grained Kullback--Leibler Control of Diffusion-Based Generative AI
by: Tsuruyama, Tatsuaki
Published: (2026)
by: Tsuruyama, Tatsuaki
Published: (2026)
Orthogonal Nonnegative Matrix Factorization with the Kullback-Leibler divergence
by: Nkurunziza, Jean Pacifique, et al.
Published: (2024)
by: Nkurunziza, Jean Pacifique, et al.
Published: (2024)
Jeffrey's update rule as a minimizer of Kullback-Leibler divergence
by: Pinzón, Carlos, et al.
Published: (2025)
by: Pinzón, Carlos, et al.
Published: (2025)
Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization
by: Jun, Kwang-Sung, et al.
Published: (2024)
by: Jun, Kwang-Sung, et al.
Published: (2024)
ScoreFusion: Fusing Score-based Generative Models via Kullback-Leibler Barycenters
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
by: Ryu, J. Jon, et al.
Published: (2025)
by: Ryu, J. Jon, et al.
Published: (2025)
Multi-agent Multi-armed Bandits with Minimum Reward Guarantee Fairness
by: Manupriya, Piyushi, et al.
Published: (2025)
by: Manupriya, Piyushi, et al.
Published: (2025)
Nearly Minimax Discrete Distribution Estimation in Kullback-Leibler Divergence with High Probability
by: van der Hoeven, Dirk, et al.
Published: (2025)
by: van der Hoeven, Dirk, et al.
Published: (2025)
Relaxed Triangle Inequality for Kullback-Leibler Divergence Between Multivariate Gaussian Distributions
by: Xiao, Shiji, et al.
Published: (2026)
by: Xiao, Shiji, et al.
Published: (2026)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
by: Amini, Afra, et al.
Published: (2025)
by: Amini, Afra, et al.
Published: (2025)
Statistical and Geometrical properties of regularized Kernel Kullback-Leibler divergence
by: Chazal, Clémentine, et al.
Published: (2024)
by: Chazal, Clémentine, et al.
Published: (2024)
CAKD: A Correlation-Aware Knowledge Distillation Framework Based on Decoupling Kullback-Leibler Divergence
by: Zhang, Zao, et al.
Published: (2024)
by: Zhang, Zao, et al.
Published: (2024)
Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge Distillation
by: Lv, Jiaming, et al.
Published: (2024)
by: Lv, Jiaming, et al.
Published: (2024)
Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
Limits to classification performance by relating Kullback-Leibler divergence to Cohen's Kappa
by: Crow, L., et al.
Published: (2024)
by: Crow, L., et al.
Published: (2024)
Entropy and the Kullback-Leibler Divergence for Bayesian Networks: Computational Complexity and Efficient Implementation
by: Scutari, Marco
Published: (2023)
by: Scutari, Marco
Published: (2023)
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss
by: Li, Yinan, et al.
Published: (2025)
by: Li, Yinan, et al.
Published: (2025)
Multi-armed Bandits with Missing Outcome
by: Mahrooghi, Ilia, et al.
Published: (2024)
by: Mahrooghi, Ilia, et al.
Published: (2024)
A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits
by: Lee, Junghyun, et al.
Published: (2024)
by: Lee, Junghyun, et al.
Published: (2024)
Kullback-Leibler excess risk bounds for exponential weighted aggregation in Generalized linear models
by: Mai, The Tien
Published: (2025)
by: Mai, The Tien
Published: (2025)
The Jacobian and Hessian of the Kullback-Leibler Divergence between Multivariate Gaussian Distributions (Technical Report)
by: Maroñas, Juan
Published: (2025)
by: Maroñas, Juan
Published: (2025)
Towards Fundamental Limits for Active Multi-distribution Learning
by: Zhang, Chicheng, et al.
Published: (2025)
by: Zhang, Chicheng, et al.
Published: (2025)
A Tight Lower Bound for Non-stochastic Multi-armed Bandits with Expert Advice
by: Chase, Zachary, et al.
Published: (2025)
by: Chase, Zachary, et al.
Published: (2025)
Probabilistic classification from possibilistic data: computing Kullback-Leibler projection with a possibility distribution
by: Baaj, Ismaïl, et al.
Published: (2026)
by: Baaj, Ismaïl, et al.
Published: (2026)
Similar Items
-
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
by: Qin, Hao, et al.
Published: (2025) -
Efficient Low-Rank Matrix Estimation, Experimental Design, and Arm-Set-Dependent Low-Rank Bandits
by: Jang, Kyoungseok, et al.
Published: (2024) -
Taming the Monster Every Context: Complexity Measure and Unified Framework for Offline-Oracle Efficient Contextual Bandits
by: Qin, Hao, et al.
Published: (2026) -
Transfer in Sequential Multi-armed Bandits via Reward Samples
by: R, Rahul N, et al.
Published: (2024) -
Decoupled Kullback-Leibler Divergence Loss
by: Cui, Jiequan, et al.
Published: (2023)