Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Schlisselberg, Ofir, Lancewicki, Tal, Auer, Peter, Mansour, Yishay
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2505.24193
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912658733662208
author Schlisselberg, Ofir
Lancewicki, Tal
Auer, Peter
Mansour, Yishay
author_facet Schlisselberg, Ofir
Lancewicki, Tal
Auer, Peter
Mansour, Yishay
contents We study the multi-armed bandit problem with adversarially chosen delays in the Best-of-Both-Worlds (BoBW) framework, which aims to achieve near-optimal performance in both stochastic and adversarial environments. While prior work has made progress toward this goal, existing algorithms suffer from significant gaps to the known lower bounds, especially in the stochastic settings. Our main contribution is a new algorithm that, up to logarithmic factors, matches the known lower bounds in each setting individually. In the adversarial case, our algorithm achieves regret of $\widetilde{O}(\sqrt{KT} + \sqrt{D})$, which is optimal up to logarithmic terms, where $T$ is the number of rounds, $K$ is the number of arms, and $D$ is the cumulative delay. In the stochastic case, we provide a regret bound which scale as $\sum_{i:Δ_i>0}\left(\log T/Δ_i\right) + \frac{1}{K}\sum Δ_i σ_{max}$, where $Δ_i$ is the sub-optimality gap of arm $i$ and $σ_{\max}$ is the maximum number of missing observations. To the best of our knowledge, this is the first BoBW algorithm to simultaneously match the lower bounds in both stochastic and adversarial regimes in delayed environment. Moreover, even beyond the BoBW setting, our stochastic regret bound is the first to match the known lower bound under adversarial delays, improving the second term over the best known result by a factor of $K$.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
Schlisselberg, Ofir
Lancewicki, Tal
Auer, Peter
Mansour, Yishay
Machine Learning
We study the multi-armed bandit problem with adversarially chosen delays in the Best-of-Both-Worlds (BoBW) framework, which aims to achieve near-optimal performance in both stochastic and adversarial environments. While prior work has made progress toward this goal, existing algorithms suffer from significant gaps to the known lower bounds, especially in the stochastic settings. Our main contribution is a new algorithm that, up to logarithmic factors, matches the known lower bounds in each setting individually. In the adversarial case, our algorithm achieves regret of $\widetilde{O}(\sqrt{KT} + \sqrt{D})$, which is optimal up to logarithmic terms, where $T$ is the number of rounds, $K$ is the number of arms, and $D$ is the cumulative delay. In the stochastic case, we provide a regret bound which scale as $\sum_{i:Δ_i>0}\left(\log T/Δ_i\right) + \frac{1}{K}\sum Δ_i σ_{max}$, where $Δ_i$ is the sub-optimality gap of arm $i$ and $σ_{\max}$ is the maximum number of missing observations. To the best of our knowledge, this is the first BoBW algorithm to simultaneously match the lower bounds in both stochastic and adversarial regimes in delayed environment. Moreover, even beyond the BoBW setting, our stochastic regret bound is the first to match the known lower bound under adversarial delays, improving the second term over the best known result by a factor of $K$.
title Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback
topic Machine Learning
url https://arxiv.org/abs/2505.24193