Zero-Inflated Bandits

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wei, Haoyu, Wan, Runzhe, Shi, Lei, Song, Rui
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917907846397952
author Wei, Haoyu
Wan, Runzhe
Shi, Lei
Song, Rui
author_facet Wei, Haoyu
Wan, Runzhe
Shi, Lei
Song, Rui
contents Many real-world bandit applications are characterized by sparse rewards, which can significantly hinder learning efficiency. Leveraging problem-specific structures for careful distribution modeling is recognized as essential for improving estimation efficiency in statistics. However, this approach remains under-explored in the context of bandits. To address this gap, we initiate the study of zero-inflated bandits, where the reward is modeled using a classic semi-parametric distribution known as the zero-inflated distribution. We develop algorithms based on the Upper Confidence Bound and Thompson Sampling frameworks for this specific structure. The superior empirical performance of these methods is demonstrated through extensive numerical studies.
format Preprint
id arxiv_https___arxiv_org_abs_2312_15595
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Zero-Inflated Bandits
Wei, Haoyu
Wan, Runzhe
Shi, Lei
Song, Rui
Machine Learning
Econometrics
Many real-world bandit applications are characterized by sparse rewards, which can significantly hinder learning efficiency. Leveraging problem-specific structures for careful distribution modeling is recognized as essential for improving estimation efficiency in statistics. However, this approach remains under-explored in the context of bandits. To address this gap, we initiate the study of zero-inflated bandits, where the reward is modeled using a classic semi-parametric distribution known as the zero-inflated distribution. We develop algorithms based on the Upper Confidence Bound and Thompson Sampling frameworks for this specific structure. The superior empirical performance of these methods is demonstrated through extensive numerical studies.
title Zero-Inflated Bandits
topic Machine Learning
Econometrics
url https://arxiv.org/abs/2312.15595