Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gadot, Uri, Derman, Esther, Kumar, Navdeep, Elfatihi, Maxence Mohamed, Levy, Kfir, Mannor, Shie
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914673163501568
author Gadot, Uri
Derman, Esther
Kumar, Navdeep
Elfatihi, Maxence Mohamed
Levy, Kfir
Mannor, Shie
author_facet Gadot, Uri
Derman, Esther
Kumar, Navdeep
Elfatihi, Maxence Mohamed
Levy, Kfir
Mannor, Shie
contents In robust Markov decision processes (RMDPs), it is assumed that the reward and the transition dynamics lie in a given uncertainty set. By targeting maximal return under the most adversarial model from that set, RMDPs address performance sensitivity to misspecified environments. Yet, to preserve computational tractability, the uncertainty set is traditionally independently structured for each state. This so-called rectangularity condition is solely motivated by computational concerns. As a result, it lacks a practical incentive and may lead to overly conservative behavior. In this work, we study coupled reward RMDPs where the transition kernel is fixed, but the reward function lies within an $α$-radius from a nominal one. We draw a direct connection between this type of non-rectangular reward-RMDPs and applying policy visitation frequency regularization. We introduce a policy-gradient method and prove its convergence. Numerical experiments illustrate the learned policy's robustness and its less conservative behavior when compared to rectangular uncertainty.
format Preprint
id arxiv_https___arxiv_org_abs_2309_01107
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
Gadot, Uri
Derman, Esther
Kumar, Navdeep
Elfatihi, Maxence Mohamed
Levy, Kfir
Mannor, Shie
Machine Learning
In robust Markov decision processes (RMDPs), it is assumed that the reward and the transition dynamics lie in a given uncertainty set. By targeting maximal return under the most adversarial model from that set, RMDPs address performance sensitivity to misspecified environments. Yet, to preserve computational tractability, the uncertainty set is traditionally independently structured for each state. This so-called rectangularity condition is solely motivated by computational concerns. As a result, it lacks a practical incentive and may lead to overly conservative behavior. In this work, we study coupled reward RMDPs where the transition kernel is fixed, but the reward function lies within an $α$-radius from a nominal one. We draw a direct connection between this type of non-rectangular reward-RMDPs and applying policy visitation frequency regularization. We introduce a policy-gradient method and prove its convergence. Numerical experiments illustrate the learned policy's robustness and its less conservative behavior when compared to rectangular uncertainty.
title Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
topic Machine Learning
url https://arxiv.org/abs/2309.01107