The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Laixi, Li, Gen, Wei, Yuting, Chen, Yuxin, Geist, Matthieu, Chi, Yuejie
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909775084650496
author Shi, Laixi
Li, Gen
Wei, Yuting
Chen, Yuxin
Geist, Matthieu
Chi, Yuejie
author_facet Shi, Laixi
Li, Gen
Wei, Yuting
Chen, Yuxin
Geist, Matthieu
Chi, Yuejie
contents This paper investigates model robustness in reinforcement learning (RL) to reduce the sim-to-real gap in practice. We adopt the framework of distributionally robust Markov decision processes (RMDPs), aimed at learning a policy that optimizes the worst-case performance when the deployed environment falls within a prescribed uncertainty set around the nominal MDP. Despite recent efforts, the sample complexity of RMDPs remained mostly unsettled regardless of the uncertainty set in use. It was unclear if distributional robustness bears any statistical consequences when benchmarked against standard RL. Assuming access to a generative model that draws samples based on the nominal MDP, we provide a near-optimal characterization of the sample complexity of RMDPs when the uncertainty set is specified via either the total variation (TV) distance or chi-squared divergence. The algorithm studied here is a model-based method called distributionally robust value iteration, which is shown to be near-optimal for the full range of uncertainty levels. Somewhat surprisingly, our results uncover that RMDPs are not necessarily easier or harder to learn than standard MDPs. The statistical consequence incurred by the robustness requirement depends heavily on the size and shape of the uncertainty set: in the case w.r.t.~the TV distance, the minimax sample complexity of RMDPs is always smaller than that of standard MDPs; in the case w.r.t.~the chi-squared divergence, the sample complexity of RMDPs far exceeds the standard MDP counterpart.
format Preprint
id arxiv_https___arxiv_org_abs_2305_16589
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model
Shi, Laixi
Li, Gen
Wei, Yuting
Chen, Yuxin
Geist, Matthieu
Chi, Yuejie
Machine Learning
Information Theory
Statistics Theory
This paper investigates model robustness in reinforcement learning (RL) to reduce the sim-to-real gap in practice. We adopt the framework of distributionally robust Markov decision processes (RMDPs), aimed at learning a policy that optimizes the worst-case performance when the deployed environment falls within a prescribed uncertainty set around the nominal MDP. Despite recent efforts, the sample complexity of RMDPs remained mostly unsettled regardless of the uncertainty set in use. It was unclear if distributional robustness bears any statistical consequences when benchmarked against standard RL. Assuming access to a generative model that draws samples based on the nominal MDP, we provide a near-optimal characterization of the sample complexity of RMDPs when the uncertainty set is specified via either the total variation (TV) distance or chi-squared divergence. The algorithm studied here is a model-based method called distributionally robust value iteration, which is shown to be near-optimal for the full range of uncertainty levels. Somewhat surprisingly, our results uncover that RMDPs are not necessarily easier or harder to learn than standard MDPs. The statistical consequence incurred by the robustness requirement depends heavily on the size and shape of the uncertainty set: in the case w.r.t.~the TV distance, the minimax sample complexity of RMDPs is always smaller than that of standard MDPs; in the case w.r.t.~the chi-squared divergence, the sample complexity of RMDPs far exceeds the standard MDP counterpart.
title The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model
topic Machine Learning
Information Theory
Statistics Theory
url https://arxiv.org/abs/2305.16589