Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Seok-Jin, Oh, Min-hwan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929677780647936
author Kim, Seok-Jin
Oh, Min-hwan
author_facet Kim, Seok-Jin
Oh, Min-hwan
contents We study the performance guarantees of exploration-free greedy algorithms for the linear contextual bandit problem. We introduce a novel condition, named the \textit{Local Anti-Concentration} (LAC) condition, which enables a greedy bandit algorithm to achieve provable efficiency. We show that the LAC condition is satisfied by a broad class of distributions, including Gaussian, exponential, uniform, Cauchy, and Student's~$t$ distributions, along with other exponential family distributions and their truncated variants. This significantly expands the class of distributions under which greedy algorithms can perform efficiently. Under our proposed LAC condition, we prove that the cumulative expected regret of the greedy algorithm for the linear contextual bandit is bounded by $O(\operatorname{poly} \log T)$. Our results establish the widest range of distributions known to date that allow a sublinear regret bound for greedy algorithms, further achieving a sharp poly-logarithmic regret.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12878
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit
Kim, Seok-Jin
Oh, Min-hwan
Machine Learning
We study the performance guarantees of exploration-free greedy algorithms for the linear contextual bandit problem. We introduce a novel condition, named the \textit{Local Anti-Concentration} (LAC) condition, which enables a greedy bandit algorithm to achieve provable efficiency. We show that the LAC condition is satisfied by a broad class of distributions, including Gaussian, exponential, uniform, Cauchy, and Student's~$t$ distributions, along with other exponential family distributions and their truncated variants. This significantly expands the class of distributions under which greedy algorithms can perform efficiently. Under our proposed LAC condition, we prove that the cumulative expected regret of the greedy algorithm for the linear contextual bandit is bounded by $O(\operatorname{poly} \log T)$. Our results establish the widest range of distributions known to date that allow a sublinear regret bound for greedy algorithms, further achieving a sharp poly-logarithmic regret.
title Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit
topic Machine Learning
url https://arxiv.org/abs/2411.12878