Intermittently Observable Markov Decision Processes

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Gongpu, Liew, Soung-Chang
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913689921126400
author Chen, Gongpu
Liew, Soung-Chang
author_facet Chen, Gongpu
Liew, Soung-Chang
contents This paper investigates MDPs with intermittent state information. We consider a scenario where the controller perceives the state information of the process via an unreliable communication channel. The transmissions of state information over the whole time horizon are modeled as a Bernoulli lossy process. Hence, the problem is finding an optimal policy for selecting actions in the presence of state information losses. We first formulate the problem as a belief MDP to establish structural results. The effect of state information losses on the expected total discounted reward is studied systematically. Then, we reformulate the problem as a tree MDP whose state space is organized in a tree structure. Two finite-state approximations to the tree MDP are developed to find near-optimal policies efficiently. Finally, we put forth a nested value iteration algorithm for the finite-state approximations, which is proved to be faster than standard value iteration. Numerical results demonstrate the effectiveness of our methods.
format Preprint
id arxiv_https___arxiv_org_abs_2302_11761
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Intermittently Observable Markov Decision Processes
Chen, Gongpu
Liew, Soung-Chang
Artificial Intelligence
Systems and Control
This paper investigates MDPs with intermittent state information. We consider a scenario where the controller perceives the state information of the process via an unreliable communication channel. The transmissions of state information over the whole time horizon are modeled as a Bernoulli lossy process. Hence, the problem is finding an optimal policy for selecting actions in the presence of state information losses. We first formulate the problem as a belief MDP to establish structural results. The effect of state information losses on the expected total discounted reward is studied systematically. Then, we reformulate the problem as a tree MDP whose state space is organized in a tree structure. Two finite-state approximations to the tree MDP are developed to find near-optimal policies efficiently. Finally, we put forth a nested value iteration algorithm for the finite-state approximations, which is proved to be faster than standard value iteration. Numerical results demonstrate the effectiveness of our methods.
title Intermittently Observable Markov Decision Processes
topic Artificial Intelligence
Systems and Control
url https://arxiv.org/abs/2302.11761