Variance-Aware Feel-Good Thompson Sampling for Contextual Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xuheng, Gu, Quanquan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feel-Good Thompson Sampling for Contextual Dueling Bandits
by: Li, Xuheng, et al.
Published: (2024)
by: Li, Xuheng, et al.
Published: (2024)
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Understanding SGD with Exponential Moving Average: A Case Study in Linear Regression
by: Li, Xuheng, et al.
Published: (2025)
by: Li, Xuheng, et al.
Published: (2025)
Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence
by: Zhang, Shiyuan, et al.
Published: (2026)
by: Zhang, Shiyuan, et al.
Published: (2026)
MARS-M: When Variance Reduction Meets Matrices
by: Liu, Yifeng, et al.
Published: (2025)
by: Liu, Yifeng, et al.
Published: (2025)
Modified Meta-Thompson Sampling for Linear Bandits and Its Bayes Regret Analysis
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
MARS: Unleashing the Power of Variance Reduction for Training Large Models
by: Yuan, Huizhuo, et al.
Published: (2024)
by: Yuan, Huizhuo, et al.
Published: (2024)
Decentralized Contextual Bandits with Network Adaptivity
by: Deng, Chuyun, et al.
Published: (2025)
by: Deng, Chuyun, et al.
Published: (2025)
Contextual Bandits with Budgeted Information Reveal
by: Gan, Kyra, et al.
Published: (2023)
by: Gan, Kyra, et al.
Published: (2023)
Unified Convergence Analysis for Score-Based Diffusion Models with Deterministic Samplers
by: Li, Runjia, et al.
Published: (2024)
by: Li, Runjia, et al.
Published: (2024)
Signature Approach for Contextual Bandits with Nonlinear and Path-dependent Rewards
by: Guo, Xin, et al.
Published: (2026)
by: Guo, Xin, et al.
Published: (2026)
Gaussian Process Thompson Sampling via Rootfinding
by: Adebiyi, Taiwo A., et al.
Published: (2024)
by: Adebiyi, Taiwo A., et al.
Published: (2024)
Epsilon-Greedy Thompson Sampling to Bayesian Optimization
by: Do, Bach, et al.
Published: (2024)
by: Do, Bach, et al.
Published: (2024)
Multi-User Contextual Cascading Bandits for Personalized Recommendation
by: Park, Jiho, et al.
Published: (2025)
by: Park, Jiho, et al.
Published: (2025)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023)
by: Zhang, Junkai, et al.
Published: (2023)
A Nearly Optimal and Low-Switching Algorithm for Reinforcement Learning with General Function Approximation
by: Zhao, Heyang, et al.
Published: (2023)
by: Zhao, Heyang, et al.
Published: (2023)
MINTS: Minimalist Thompson Sampling
by: Wang, Kaizheng
Published: (2026)
by: Wang, Kaizheng
Published: (2026)
Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning
by: Di, Qiwei, et al.
Published: (2023)
by: Di, Qiwei, et al.
Published: (2023)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
by: Kou, Yiwen, et al.
Published: (2024)
by: Kou, Yiwen, et al.
Published: (2024)
Sample Complexity of Variance-reduced Distributionally Robust Q-learning
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
Deconfounded Warm-Start Thompson Sampling with Applications to Precision Medicine
by: Jaiswal, Prateek, et al.
Published: (2025)
by: Jaiswal, Prateek, et al.
Published: (2025)
Early Stopping in Contextual Bandits and Inferences
by: Cui, Zihan
Published: (2025)
by: Cui, Zihan
Published: (2025)
Variance-Reduced Cascade Q-learning: Algorithms and Sample Complexity
by: Boveiri, Mohammad, et al.
Published: (2024)
by: Boveiri, Mohammad, et al.
Published: (2024)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
by: Zhou, Dongruo, et al.
Published: (2018)
by: Zhou, Dongruo, et al.
Published: (2018)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
by: Zhang, Weitong, et al.
Published: (2021)
by: Zhang, Weitong, et al.
Published: (2021)
Variance-Dependent Regret Lower Bounds for Contextual Bandits
by: He, Jiafan, et al.
Published: (2025)
by: He, Jiafan, et al.
Published: (2025)
Distributed Thompson sampling under constrained communication
by: Zerefa, Saba, et al.
Published: (2024)
by: Zerefa, Saba, et al.
Published: (2024)
Towards Simple and Provable Parameter-Free Adaptive Gradient Methods
by: Tao, Yuanzhe, et al.
Published: (2024)
by: Tao, Yuanzhe, et al.
Published: (2024)
Variance Reduction and Low Sample Complexity in Stochastic Optimization via Proximal Point Method
by: Liang, Jiaming
Published: (2024)
by: Liang, Jiaming
Published: (2024)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
by: Chen, Zixiang, et al.
Published: (2025)
by: Chen, Zixiang, et al.
Published: (2025)
Frequentist Regret Analysis of Gaussian Process Thompson Sampling via Fractional Posteriors
by: Roy, Somjit, et al.
Published: (2026)
by: Roy, Somjit, et al.
Published: (2026)
Bandit Convex Optimisation
by: Lattimore, Tor
Published: (2024)
by: Lattimore, Tor
Published: (2024)
Optimism Stabilizes Thompson Sampling for Adaptive Inference
by: Yan, Shunxing, et al.
Published: (2026)
by: Yan, Shunxing, et al.
Published: (2026)
Reinforcement Learning from Human Feedback with Active Queries
by: Ji, Kaixuan, et al.
Published: (2024)
by: Ji, Kaixuan, et al.
Published: (2024)
Accelerating RLHF Training with Reward Variance Increase
by: Yang, Zonglin, et al.
Published: (2025)
by: Yang, Zonglin, et al.
Published: (2025)
Beyond Bounded Variance: Variance-Reduced Normalized Methods for Nonconvex Optimization under Blum-Gladyshev Noise
by: Upadhyay, Antesh, et al.
Published: (2026)
by: Upadhyay, Antesh, et al.
Published: (2026)
Analysis of Thompson Sampling for Controlling Unknown Linear Diffusion Processes
by: Faradonbeh, Mohamad Kazem Shirani, et al.
Published: (2022)
by: Faradonbeh, Mohamad Kazem Shirani, et al.
Published: (2022)
The Safety-Privacy Tradeoff in Linear Bandits
by: Zibaie, Arghavan, et al.
Published: (2025)
by: Zibaie, Arghavan, et al.
Published: (2025)
Muon is Provably Faster with Momentum Variance Reduction
by: Qian, Xun, et al.
Published: (2025)
by: Qian, Xun, et al.
Published: (2025)
Towards Weaker Variance Assumptions for Stochastic Optimization
by: Alacaoglu, Ahmet, et al.
Published: (2025)
by: Alacaoglu, Ahmet, et al.
Published: (2025)
Similar Items
-
Feel-Good Thompson Sampling for Contextual Dueling Bandits
by: Li, Xuheng, et al.
Published: (2024) -
Variance-Aware Regret Bounds for Stochastic Contextual Dueling Bandits
by: Di, Qiwei, et al.
Published: (2023) -
Understanding SGD with Exponential Moving Average: A Case Study in Linear Regression
by: Li, Xuheng, et al.
Published: (2025) -
Dimension-Independent Convergence of Underdamped Langevin Monte Carlo in KL Divergence
by: Zhang, Shiyuan, et al.
Published: (2026) -
MARS-M: When Variance Reduction Meets Matrices
by: Liu, Yifeng, et al.
Published: (2025)