Stochastic Bandits with ReLU Neural Networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Kan, Bastani, Hamsa, Goel, Surbhi, Bastani, Osbert
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914793601892352
author Xu, Kan
Bastani, Hamsa
Goel, Surbhi
Bastani, Osbert
author_facet Xu, Kan
Bastani, Hamsa
Goel, Surbhi
Bastani, Osbert
contents We study the stochastic bandit problem with ReLU neural network structure. We show that a $\tilde{O}(\sqrt{T})$ regret guarantee is achievable by considering bandits with one-layer ReLU neural networks; to the best of our knowledge, our work is the first to achieve such a guarantee. In this specific setting, we propose an OFU-ReLU algorithm that can achieve this upper bound. The algorithm first explores randomly until it reaches a linear regime, and then implements a UCB-type linear bandit algorithm to balance exploration and exploitation. Our key insight is that we can exploit the piecewise linear structure of ReLU activations and convert the problem into a linear bandit in a transformed feature space, once we learn the parameters of ReLU relatively accurately during the exploration stage. To remove dependence on model parameters, we design an OFU-ReLU+ algorithm based on a batching strategy, which can provide the same theoretical guarantee.
format Preprint
id arxiv_https___arxiv_org_abs_2405_07331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stochastic Bandits with ReLU Neural Networks
Xu, Kan
Bastani, Hamsa
Goel, Surbhi
Bastani, Osbert
Machine Learning
Data Structures and Algorithms
We study the stochastic bandit problem with ReLU neural network structure. We show that a $\tilde{O}(\sqrt{T})$ regret guarantee is achievable by considering bandits with one-layer ReLU neural networks; to the best of our knowledge, our work is the first to achieve such a guarantee. In this specific setting, we propose an OFU-ReLU algorithm that can achieve this upper bound. The algorithm first explores randomly until it reaches a linear regime, and then implements a UCB-type linear bandit algorithm to balance exploration and exploitation. Our key insight is that we can exploit the piecewise linear structure of ReLU activations and convert the problem into a linear bandit in a transformed feature space, once we learn the parameters of ReLU relatively accurately during the exploration stage. To remove dependence on model parameters, we design an OFU-ReLU+ algorithm based on a batching strategy, which can provide the same theoretical guarantee.
title Stochastic Bandits with ReLU Neural Networks
topic Machine Learning
Data Structures and Algorithms
url https://arxiv.org/abs/2405.07331