Online learning in bandits with predicted context
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Yongyi, Xu, Ziping, Murphy, Susan |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Statistical Reinforcement Learning in the Real World: A Survey of Challenges and Future Directions
by: Gazi, Asim H., et al.
Published: (2026)
by: Gazi, Asim H., et al.
Published: (2026)
Optimal cross-learning for contextual bandits with unknown context distributions
by: Schneider, Jon, et al.
Published: (2024)
by: Schneider, Jon, et al.
Published: (2024)
The Fallacy of Minimizing Cumulative Regret in the Sequential Task Setting
by: Xu, Ziping, et al.
Published: (2024)
by: Xu, Ziping, et al.
Published: (2024)
A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage
by: Tan, Kevin, et al.
Published: (2024)
by: Tan, Kevin, et al.
Published: (2024)
Active Measuring in Reinforcement Learning With Delayed Negative Effects
by: Gao, Daiqi, et al.
Published: (2025)
by: Gao, Daiqi, et al.
Published: (2025)
Reinforcement learning with combinatorial actions for coupled restless bandits
by: Xu, Lily, et al.
Published: (2025)
by: Xu, Lily, et al.
Published: (2025)
Truthful mechanisms for linear bandit games with private contexts
by: Hu, Yiting, et al.
Published: (2025)
by: Hu, Yiting, et al.
Published: (2025)
Extreme bandits
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
reBandit: Random Effects based Online RL algorithm for Reducing Cannabis Use
by: Ghosh, Susobhan, et al.
Published: (2024)
by: Ghosh, Susobhan, et al.
Published: (2024)
Model selection for behavioral learning data and applications to contextual bandits
by: Aubert, Julien, et al.
Published: (2025)
by: Aubert, Julien, et al.
Published: (2025)
Efficient learning by implicit exploration in bandit problems with side observations
by: Kocak, Tomas, et al.
Published: (2026)
by: Kocak, Tomas, et al.
Published: (2026)
Online Conformal Prediction with Adversarial Semi-bandit Feedback via Regret Minimization
by: Yang, Junyoung, et al.
Published: (2026)
by: Yang, Junyoung, et al.
Published: (2026)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Active clustering with bandit feedback
by: Thuot, Victor, et al.
Published: (2024)
by: Thuot, Victor, et al.
Published: (2024)
In-context learning to predict critical transitions in dynamical systems
by: Sevinchan, Yunus, et al.
Published: (2026)
by: Sevinchan, Yunus, et al.
Published: (2026)
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
by: Mei, Jincheng, et al.
Published: (2025)
by: Mei, Jincheng, et al.
Published: (2025)
Approximate information maximization for bandit games
by: Barbier-Chebbah, Alex, et al.
Published: (2023)
by: Barbier-Chebbah, Alex, et al.
Published: (2023)
Instance-dependent Stochastic Lipschitz bandit
by: Potfer, Marius, et al.
Published: (2026)
by: Potfer, Marius, et al.
Published: (2026)
Spectral bandits for smooth graph functions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
From predictions to confidence intervals: an empirical study of conformal prediction methods for in-context learning
by: Huang, Zhe, et al.
Published: (2025)
by: Huang, Zhe, et al.
Published: (2025)
Missing Data Multiple Imputation for Tabular Q-Learning in Online RL
by: Chasalow, Kyla, et al.
Published: (2025)
by: Chasalow, Kyla, et al.
Published: (2025)
Non-Stationary Latent Auto-Regressive Bandits
by: Trella, Anna L., et al.
Published: (2024)
by: Trella, Anna L., et al.
Published: (2024)
Statistical Inference for Misspecified Contextual Bandits
by: Guo, Yongyi, et al.
Published: (2025)
by: Guo, Yongyi, et al.
Published: (2025)
Risk and optimal policies in bandit experiments
by: Adusumilli, Karun
Published: (2021)
by: Adusumilli, Karun
Published: (2021)
Revealing graph bandits for maximizing local influence
by: Carpentier, Alexandra, et al.
Published: (2026)
by: Carpentier, Alexandra, et al.
Published: (2026)
On the optimal regret of collaborative personalized linear bandits
by: Huang, Bruce, et al.
Published: (2025)
by: Huang, Bruce, et al.
Published: (2025)
Offline-to-online hyperparameter transfer for stochastic bandits
by: Sharma, Dravyansh, et al.
Published: (2025)
by: Sharma, Dravyansh, et al.
Published: (2025)
Linear bandits with polylogarithmic minimax regret
by: Lumbreras, Josep, et al.
Published: (2024)
by: Lumbreras, Josep, et al.
Published: (2024)
Reinforcement Learning on Dyads to Enhance Medication Adherence
by: Xu, Ziping, et al.
Published: (2025)
by: Xu, Ziping, et al.
Published: (2025)
Ensemble sampling for linear bandits: small ensembles suffice
by: Janz, David, et al.
Published: (2023)
by: Janz, David, et al.
Published: (2023)
VITS : Variational Inference Thompson Sampling for contextual bandits
by: Clavier, Pierre, et al.
Published: (2023)
by: Clavier, Pierre, et al.
Published: (2023)
Lookahead identification in adversarial bandits: accuracy and memory bounds
by: Brukhim, Nataly, et al.
Published: (2026)
by: Brukhim, Nataly, et al.
Published: (2026)
Efficient kernelized bandit algorithms via exploration distributions
by: Hu, Bingshan, et al.
Published: (2025)
by: Hu, Bingshan, et al.
Published: (2025)
Leveraging priors on distribution functions for multi-arm bandits
by: Vashishtha, Sumit, et al.
Published: (2025)
by: Vashishtha, Sumit, et al.
Published: (2025)
Trading off rewards and errors in multi-armed bandits
by: Erraqabi, Akram, et al.
Published: (2026)
by: Erraqabi, Akram, et al.
Published: (2026)
Minimum mean-squared error estimation with bandit feedback
by: Ghosh, Ayon, et al.
Published: (2022)
by: Ghosh, Ayon, et al.
Published: (2022)
When and why randomised exploration works (in linear bandits)
by: Abeille, Marc, et al.
Published: (2025)
by: Abeille, Marc, et al.
Published: (2025)
Learning with Incomplete Context: Linear Contextual Bandits with Pretrained Imputation
by: Yan, Hao, et al.
Published: (2025)
by: Yan, Hao, et al.
Published: (2025)
Online Uniform Sampling: Randomized Learning-Augmented Approximation Algorithms with Application to Digital Health
by: Liu, Xueqing, et al.
Published: (2024)
by: Liu, Xueqing, et al.
Published: (2024)
Best-of-Both Worlds for linear contextual bandits with paid observations
by: Boyer, Nathan, et al.
Published: (2025)
by: Boyer, Nathan, et al.
Published: (2025)
Similar Items
-
Statistical Reinforcement Learning in the Real World: A Survey of Challenges and Future Directions
by: Gazi, Asim H., et al.
Published: (2026) -
Optimal cross-learning for contextual bandits with unknown context distributions
by: Schneider, Jon, et al.
Published: (2024) -
The Fallacy of Minimizing Cumulative Regret in the Sequential Task Setting
by: Xu, Ziping, et al.
Published: (2024) -
A Natural Extension To Online Algorithms For Hybrid RL With Limited Coverage
by: Tan, Kevin, et al.
Published: (2024) -
Active Measuring in Reinforcement Learning With Delayed Negative Effects
by: Gao, Daiqi, et al.
Published: (2025)