Incentivized Exploration of Non-Stationary Stochastic Bandits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chakraborty, Sourav, Chen, Lijun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911799274635264
author Chakraborty, Sourav
Chen, Lijun
author_facet Chakraborty, Sourav
Chen, Lijun
contents We study incentivized exploration for the multi-armed bandit (MAB) problem with non-stationary reward distributions, where players receive compensation for exploring arms other than the greedy choice and may provide biased feedback on the reward. We consider two different non-stationary environments: abruptly-changing and continuously-changing, and propose respective incentivized exploration algorithms. We show that the proposed algorithms achieve sublinear regret and compensation over time, thus effectively incentivizing exploration despite the nonstationarity and the biased or drifted feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10819
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Incentivized Exploration of Non-Stationary Stochastic Bandits
Chakraborty, Sourav
Chen, Lijun
Machine Learning
Artificial Intelligence
We study incentivized exploration for the multi-armed bandit (MAB) problem with non-stationary reward distributions, where players receive compensation for exploring arms other than the greedy choice and may provide biased feedback on the reward. We consider two different non-stationary environments: abruptly-changing and continuously-changing, and propose respective incentivized exploration algorithms. We show that the proposed algorithms achieve sublinear regret and compensation over time, thus effectively incentivizing exploration despite the nonstationarity and the biased or drifted feedback.
title Incentivized Exploration of Non-Stationary Stochastic Bandits
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2403.10819