Saved in:
Bibliographic Details
Main Authors: Liu, Xiao-Yin, Zhou, Xiao-Hu, Gui, Mei-Jiang, Li, Guo-Tao, Xie, Xiao-Liang, Liu, Shi-Qi, Wang, Shuang-Yi, Zhang, Qi-Chao, Luo, Biao, Hou, Zeng-Guang
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2309.08925
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916783100788736
author Liu, Xiao-Yin
Zhou, Xiao-Hu
Gui, Mei-Jiang
Li, Guo-Tao
Xie, Xiao-Liang
Liu, Shi-Qi
Wang, Shuang-Yi
Zhang, Qi-Chao
Luo, Biao
Hou, Zeng-Guang
author_facet Liu, Xiao-Yin
Zhou, Xiao-Hu
Gui, Mei-Jiang
Li, Guo-Tao
Xie, Xiao-Liang
Liu, Shi-Qi
Wang, Shuang-Yi
Zhang, Qi-Chao
Luo, Biao
Hou, Zeng-Guang
contents Model-based reinforcement learning (RL), which learns an environment model from the offline dataset and generates more out-of-distribution model data, has become an effective approach to the problem of distribution shift in offline RL. Due to the gap between the learned and actual environment, conservatism should be incorporated into the algorithm to balance accurate offline data and imprecise model data. The conservatism of current algorithms mostly relies on model uncertainty estimation. However, uncertainty estimation is unreliable and leads to poor performance in certain scenarios, and the previous methods ignore differences between the model data, which brings great conservatism. To address the above issues, this paper proposes a milDly cOnservative Model-bAsed offlINe RL algorithm (DOMAIN) without estimating model uncertainty, and designs the adaptive sampling distribution of model samples, which can adaptively adjust the model data penalty. In this paper, we theoretically demonstrate that the Q value learned by the DOMAIN outside the region is a lower bound of the true Q value, the DOMAIN is less conservative than previous model-based offline RL algorithms, and has the guarantee of safety policy improvement. The results of extensive experiments show that DOMAIN outperforms prior RL algorithms and the average performance has improved by 1.8% on the D4RL benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2309_08925
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DOMAIN: MilDly COnservative Model-BAsed OfflINe Reinforcement Learning
Liu, Xiao-Yin
Zhou, Xiao-Hu
Gui, Mei-Jiang
Li, Guo-Tao
Xie, Xiao-Liang
Liu, Shi-Qi
Wang, Shuang-Yi
Zhang, Qi-Chao
Luo, Biao
Hou, Zeng-Guang
Machine Learning
Artificial Intelligence
Model-based reinforcement learning (RL), which learns an environment model from the offline dataset and generates more out-of-distribution model data, has become an effective approach to the problem of distribution shift in offline RL. Due to the gap between the learned and actual environment, conservatism should be incorporated into the algorithm to balance accurate offline data and imprecise model data. The conservatism of current algorithms mostly relies on model uncertainty estimation. However, uncertainty estimation is unreliable and leads to poor performance in certain scenarios, and the previous methods ignore differences between the model data, which brings great conservatism. To address the above issues, this paper proposes a milDly cOnservative Model-bAsed offlINe RL algorithm (DOMAIN) without estimating model uncertainty, and designs the adaptive sampling distribution of model samples, which can adaptively adjust the model data penalty. In this paper, we theoretically demonstrate that the Q value learned by the DOMAIN outside the region is a lower bound of the true Q value, the DOMAIN is less conservative than previous model-based offline RL algorithms, and has the guarantee of safety policy improvement. The results of extensive experiments show that DOMAIN outperforms prior RL algorithms and the average performance has improved by 1.8% on the D4RL benchmark.
title DOMAIN: MilDly COnservative Model-BAsed OfflINe Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2309.08925