Stochastic Graph Bandit Learning with Side-Observations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gong, Xueping, Zhang, Jiheng
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914631454294016
author Gong, Xueping
Zhang, Jiheng
author_facet Gong, Xueping
Zhang, Jiheng
contents In this paper, we investigate the stochastic contextual bandit with general function space and graph feedback. We propose an algorithm that addresses this problem by adapting to both the underlying graph structures and reward gaps. To the best of our knowledge, our algorithm is the first to provide a gap-dependent upper bound in this stochastic setting, bridging the research gap left by the work in [35]. In comparison to [31,33,35], our method offers improved regret upper bounds and does not require knowledge of graphical quantities. We conduct numerical experiments to demonstrate the computational efficiency and effectiveness of our approach in terms of regret upper bounds. These findings highlight the significance of our algorithm in advancing the field of stochastic contextual bandits with graph feedback, opening up avenues for practical applications in various domains.
format Preprint
id arxiv_https___arxiv_org_abs_2308_15107
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Stochastic Graph Bandit Learning with Side-Observations
Gong, Xueping
Zhang, Jiheng
Machine Learning
In this paper, we investigate the stochastic contextual bandit with general function space and graph feedback. We propose an algorithm that addresses this problem by adapting to both the underlying graph structures and reward gaps. To the best of our knowledge, our algorithm is the first to provide a gap-dependent upper bound in this stochastic setting, bridging the research gap left by the work in [35]. In comparison to [31,33,35], our method offers improved regret upper bounds and does not require knowledge of graphical quantities. We conduct numerical experiments to demonstrate the computational efficiency and effectiveness of our approach in terms of regret upper bounds. These findings highlight the significance of our algorithm in advancing the field of stochastic contextual bandits with graph feedback, opening up avenues for practical applications in various domains.
title Stochastic Graph Bandit Learning with Side-Observations
topic Machine Learning
url https://arxiv.org/abs/2308.15107