Single-Trajectory Distributionally Robust Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Zhipeng, Ma, Xiaoteng, Blanchet, Jose, Zhang, Jiheng, Zhou, Zhengyuan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914953962717184
author Liang, Zhipeng
Ma, Xiaoteng
Blanchet, Jose
Zhang, Jiheng
Zhou, Zhengyuan
author_facet Liang, Zhipeng
Ma, Xiaoteng
Blanchet, Jose
Zhang, Jiheng
Zhou, Zhengyuan
contents To mitigate the limitation that the classical reinforcement learning (RL) framework heavily relies on identical training and test environments, Distributionally Robust RL (DRRL) has been proposed to enhance performance across a range of environments, possibly including unknown test environments. As a price for robustness gain, DRRL involves optimizing over a set of distributions, which is inherently more challenging than optimizing over a fixed distribution in the non-robust case. Existing DRRL algorithms are either model-based or fail to learn from a single sample trajectory. In this paper, we design a first fully model-free DRRL algorithm, called distributionally robust Q-learning with single trajectory (DRQ). We delicately design a multi-timescale framework to fully utilize each incrementally arriving sample and directly learn the optimal distributionally robust policy without modelling the environment, thus the algorithm can be trained along a single trajectory in a model-free fashion. Despite the algorithm's complexity, we provide asymptotic convergence guarantees by generalizing classical stochastic approximation tools. Comprehensive experimental results demonstrate the superior robustness and sample complexity of our proposed algorithm, compared to non-robust methods and other robust RL algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2301_11721
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Single-Trajectory Distributionally Robust Reinforcement Learning
Liang, Zhipeng
Ma, Xiaoteng
Blanchet, Jose
Zhang, Jiheng
Zhou, Zhengyuan
Machine Learning
Artificial Intelligence
To mitigate the limitation that the classical reinforcement learning (RL) framework heavily relies on identical training and test environments, Distributionally Robust RL (DRRL) has been proposed to enhance performance across a range of environments, possibly including unknown test environments. As a price for robustness gain, DRRL involves optimizing over a set of distributions, which is inherently more challenging than optimizing over a fixed distribution in the non-robust case. Existing DRRL algorithms are either model-based or fail to learn from a single sample trajectory. In this paper, we design a first fully model-free DRRL algorithm, called distributionally robust Q-learning with single trajectory (DRQ). We delicately design a multi-timescale framework to fully utilize each incrementally arriving sample and directly learn the optimal distributionally robust policy without modelling the environment, thus the algorithm can be trained along a single trajectory in a model-free fashion. Despite the algorithm's complexity, we provide asymptotic convergence guarantees by generalizing classical stochastic approximation tools. Comprehensive experimental results demonstrate the superior robustness and sample complexity of our proposed algorithm, compared to non-robust methods and other robust RL algorithms.
title Single-Trajectory Distributionally Robust Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2301.11721