Relaxed Indexability and Index Policy for Partially Observable Restless Bandits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Liu, Keqin
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915246480818176
author Liu, Keqin
author_facet Liu, Keqin
contents This paper addresses an important class of restless multi-armed bandit (RMAB) problems that finds broad application in operations research, stochastic optimization, and reinforcement learning. There are $N$ independent Markov processes that may be operated, observed and offer rewards. Due to the resource constraint, we can only choose a subset of $M~(M<N)$ processes to operate and accrue reward determined by the states of selected processes. We formulate the problem as a partially observable RMAB with an infinite state space and design an algorithm that achieves a near-optimal performance with low complexity. Our algorithm is based on a generalization of Whittle's original idea of indexability. Referred to as the relaxed indexability, the extended definition leads to the efficient online verifications and computations of the approximate Whittle index under the proposed algorithmic framework.
format Preprint
id arxiv_https___arxiv_org_abs_2107_11939
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Relaxed Indexability and Index Policy for Partially Observable Restless Bandits
Liu, Keqin
Optimization and Control
This paper addresses an important class of restless multi-armed bandit (RMAB) problems that finds broad application in operations research, stochastic optimization, and reinforcement learning. There are $N$ independent Markov processes that may be operated, observed and offer rewards. Due to the resource constraint, we can only choose a subset of $M~(M<N)$ processes to operate and accrue reward determined by the states of selected processes. We formulate the problem as a partially observable RMAB with an infinite state space and design an algorithm that achieves a near-optimal performance with low complexity. Our algorithm is based on a generalization of Whittle's original idea of indexability. Referred to as the relaxed indexability, the extended definition leads to the efficient online verifications and computations of the approximate Whittle index under the proposed algorithmic framework.
title Relaxed Indexability and Index Policy for Partially Observable Restless Bandits
topic Optimization and Control
url https://arxiv.org/abs/2107.11939