A Tight Lower Bound for Non-stochastic Multi-armed Bandits with Expert Advice
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866915589732171776 |
|---|---|
| author | Chase, Zachary Ito, Shinji Mehalel, Idan |
| author_facet | Chase, Zachary Ito, Shinji Mehalel, Idan |
| contents | We determine the minimax optimal expected regret in the classic non-stochastic multi-armed bandit with expert advice problem, by proving a lower bound that matches the upper bound of Kale (2014). The two bounds determine the minimax optimal expected regret to be $Θ\left( \sqrt{T K \log (N/K) } \right)$, where $K$ is the number of arms, $N$ is the number of experts, and $T$ is the time horizon. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_00257 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Tight Lower Bound for Non-stochastic Multi-armed Bandits with Expert Advice Chase, Zachary Ito, Shinji Mehalel, Idan Machine Learning We determine the minimax optimal expected regret in the classic non-stochastic multi-armed bandit with expert advice problem, by proving a lower bound that matches the upper bound of Kale (2014). The two bounds determine the minimax optimal expected regret to be $Θ\left( \sqrt{T K \log (N/K) } \right)$, where $K$ is the number of arms, $N$ is the number of experts, and $T$ is the time horizon. |
| title | A Tight Lower Bound for Non-stochastic Multi-armed Bandits with Expert Advice |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2511.00257 |