Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kash, Ian A., Reyzin, Lev, Yu, Zishun
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911791886368768
author Kash, Ian A.
Reyzin, Lev
Yu, Zishun
author_facet Kash, Ian A.
Reyzin, Lev
Yu, Zishun
contents Reinforcement learning generalizes multi-armed bandit problems with additional difficulties of a longer planning horizon and unknown transition kernel. We explore a black-box reduction from discounted infinite-horizon tabular reinforcement learning to multi-armed bandits, where, specifically, an independent bandit learner is placed in each state. We show that, under ergodicity and fast mixing assumptions, any slowly changing adversarial bandit algorithm achieving optimal regret in the adversarial bandit setting can also attain optimal expected regret in infinite-horizon discounted Markov decision processes, with respect to the number of rounds $T$. Furthermore, we examine our reduction using a specific instance of the exponential-weight algorithm.
format Preprint
id arxiv_https___arxiv_org_abs_2205_09056
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
Kash, Ian A.
Reyzin, Lev
Yu, Zishun
Machine Learning
Reinforcement learning generalizes multi-armed bandit problems with additional difficulties of a longer planning horizon and unknown transition kernel. We explore a black-box reduction from discounted infinite-horizon tabular reinforcement learning to multi-armed bandits, where, specifically, an independent bandit learner is placed in each state. We show that, under ergodicity and fast mixing assumptions, any slowly changing adversarial bandit algorithm achieving optimal regret in the adversarial bandit setting can also attain optimal expected regret in infinite-horizon discounted Markov decision processes, with respect to the number of rounds $T$. Furthermore, we examine our reduction using a specific instance of the exponential-weight algorithm.
title Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
topic Machine Learning
url https://arxiv.org/abs/2205.09056