On the Benefits of Free Exploration for Regret Minimization in Multi-Armed Bandits
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Yunlong, Zhong, Zixin, Tan, Vincent Y. F. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits
by: Hou, Yunlong, et al.
Published: (2024)
by: Hou, Yunlong, et al.
Published: (2024)
Asymptotically and Minimax Optimal Regret Bounds for Multi-Armed Bandits with Abstention
by: Yang, Junwen, et al.
Published: (2024)
by: Yang, Junwen, et al.
Published: (2024)
Optimal Clustering with Bandit Feedback
by: Yang, Junwen, et al.
Published: (2022)
by: Yang, Junwen, et al.
Published: (2022)
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Best Arm Identification with Minimal Regret
by: Yang, Junwen, et al.
Published: (2024)
by: Yang, Junwen, et al.
Published: (2024)
Asymptotically Optimal Linear Best Feasible Arm Identification with Fixed Budget
by: Bian, Jie, et al.
Published: (2025)
by: Bian, Jie, et al.
Published: (2025)
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
by: Chen, Zhirui, et al.
Published: (2025)
by: Chen, Zhirui, et al.
Published: (2025)
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
by: Chen, Zhirui, et al.
Published: (2024)
by: Chen, Zhirui, et al.
Published: (2024)
Best Arm Identification with Possibly Biased Offline Data
by: Yang, Le, et al.
Published: (2025)
by: Yang, Le, et al.
Published: (2025)
Diminishing Exploration: A Minimalist Approach to Piecewise Stationary Multi-Armed Bandits
by: Li, Kuan-Ta, et al.
Published: (2024)
by: Li, Kuan-Ta, et al.
Published: (2024)
Influence Maximization via Graph Neural Bandits
by: Feng, Yuting, et al.
Published: (2024)
by: Feng, Yuting, et al.
Published: (2024)
Combinatorial Multi-armed Bandits: Arm Selection via Group Testing
by: Mukherjee, Arpan, et al.
Published: (2024)
by: Mukherjee, Arpan, et al.
Published: (2024)
On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization
by: Ji, Kaixuan, et al.
Published: (2026)
by: Ji, Kaixuan, et al.
Published: (2026)
Fixed-Budget Differentially Private Best Arm Identification
by: Chen, Zhirui, et al.
Published: (2024)
by: Chen, Zhirui, et al.
Published: (2024)
Regret Bounds for Noise-Free Cascaded Kernelized Bandits
by: Li, Zihan, et al.
Published: (2022)
by: Li, Zihan, et al.
Published: (2022)
Quantile Multi-Armed Bandits with 1-bit Feedback
by: Lau, Ivan, et al.
Published: (2025)
by: Lau, Ivan, et al.
Published: (2025)
Indexed Minimum Empirical Divergence-Based Algorithms for Linear Bandits
by: Bian, Jie, et al.
Published: (2024)
by: Bian, Jie, et al.
Published: (2024)
Bandit Convex Optimization with Gradient Prediction Adaptivity
by: Wang, Shuche, et al.
Published: (2026)
by: Wang, Shuche, et al.
Published: (2026)
Flickering Multi-Armed Bandits
by: Chakraborty, Sourav, et al.
Published: (2026)
by: Chakraborty, Sourav, et al.
Published: (2026)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
by: Hou, Yunlong, et al.
Published: (2025)
by: Hou, Yunlong, et al.
Published: (2025)
ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
by: Gilani, Atefeh, et al.
Published: (2026)
by: Gilani, Atefeh, et al.
Published: (2026)
Remote Action Generation: Remote Control with Minimal Communication
by: Kobus, Szymon, et al.
Published: (2026)
by: Kobus, Szymon, et al.
Published: (2026)
Multi-Armed Bandits With Best-Action Queries
by: Bacchiocchi, Francesco, et al.
Published: (2026)
by: Bacchiocchi, Francesco, et al.
Published: (2026)
Fast Rates for Offline Contextual Bandits with Forward-KL Regularization under Single-Policy Concentrability
by: Zhao, Qingyue, et al.
Published: (2026)
by: Zhao, Qingyue, et al.
Published: (2026)
Energy Saving for Cell-Free Massive MIMO Networks: A Multi-Agent Deep Reinforcement Learning Approach
by: Wang, Qichen, et al.
Published: (2026)
by: Wang, Qichen, et al.
Published: (2026)
Evolution of Information in Interactive Decision Making: A Case Study for Multi-Armed Bandits
by: Gu, Yuzhou, et al.
Published: (2025)
by: Gu, Yuzhou, et al.
Published: (2025)
Optimal Best Arm Identification with Fixed Confidence in Restless Bandits
by: Karthik, P. N., et al.
Published: (2023)
by: Karthik, P. N., et al.
Published: (2023)
$(ε, u)$-Adaptive Regret Minimization in Heavy-Tailed Bandits
by: Genalti, Gianmarco, et al.
Published: (2023)
by: Genalti, Gianmarco, et al.
Published: (2023)
Quantum-Enhanced Neural Contextual Bandit Algorithms
by: Huang, Yuqi, et al.
Published: (2026)
by: Huang, Yuqi, et al.
Published: (2026)
Introduction to Multi-Armed Bandits
by: Slivkins, Aleksandrs
Published: (2019)
by: Slivkins, Aleksandrs
Published: (2019)
Fair Algorithms with Probing for Multi-Agent Multi-Armed Bandits
by: Xu, Tianyi, et al.
Published: (2025)
by: Xu, Tianyi, et al.
Published: (2025)
Large Language Model-Enhanced Multi-Armed Bandits
by: Sun, Jiahang, et al.
Published: (2025)
by: Sun, Jiahang, et al.
Published: (2025)
Multi-Armed Bandits-Based Optimization of Decision Trees
by: Shanto, Hasibul Karim, et al.
Published: (2025)
by: Shanto, Hasibul Karim, et al.
Published: (2025)
Intrinsic Fairness-Accuracy Tradeoffs under Equalized Odds
by: Zhong, Meiyu, et al.
Published: (2024)
by: Zhong, Meiyu, et al.
Published: (2024)
A Frequency-Domain Analysis of the Multi-Armed Bandit Problem: A New Perspective on the Exploration-Exploitation Trade-off
by: Zhang, Di
Published: (2025)
by: Zhang, Di
Published: (2025)
p-Mean Regret for Stochastic Bandits
by: Krishna, Anand, et al.
Published: (2024)
by: Krishna, Anand, et al.
Published: (2024)
Regret Tail Characterization of Optimal Bandit Algorithms with Generic Rewards
by: Panda, Subhodip, et al.
Published: (2026)
by: Panda, Subhodip, et al.
Published: (2026)
Improved Regret Bounds for Linear Bandits with Heavy-Tailed Rewards
by: Tajdini, Artin, et al.
Published: (2025)
by: Tajdini, Artin, et al.
Published: (2025)
Next-Token Prediction and Regret Minimization
by: Mohri, Mehryar, et al.
Published: (2026)
by: Mohri, Mehryar, et al.
Published: (2026)
Global Rewards in Restless Multi-Armed Bandits
by: Raman, Naveen, et al.
Published: (2024)
by: Raman, Naveen, et al.
Published: (2024)
Similar Items
-
Almost Minimax Optimal Best Arm Identification in Piecewise Stationary Linear Bandits
by: Hou, Yunlong, et al.
Published: (2024) -
Asymptotically and Minimax Optimal Regret Bounds for Multi-Armed Bandits with Abstention
by: Yang, Junwen, et al.
Published: (2024) -
Optimal Clustering with Bandit Feedback
by: Yang, Junwen, et al.
Published: (2022) -
Near-Optimal Regret for KL-Regularized Multi-Armed Bandits
by: Ji, Kaixuan, et al.
Published: (2026) -
Best Arm Identification with Minimal Regret
by: Yang, Junwen, et al.
Published: (2024)