Conservative Distributional Reinforcement Learning with Safety Constraints

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Hengrui, Lin, Youfang, Han, Sheng, Wang, Shuo, Lv, Kai
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912086278275072
author Zhang, Hengrui
Lin, Youfang
Han, Sheng
Wang, Shuo
Lv, Kai
author_facet Zhang, Hengrui
Lin, Youfang
Han, Sheng
Wang, Shuo
Lv, Kai
contents Safety exploration can be regarded as a constrained Markov decision problem where the expected long-term cost is constrained. Previous off-policy algorithms convert the constrained optimization problem into the corresponding unconstrained dual problem by introducing the Lagrangian relaxation technique. However, the cost function of the above algorithms provides inaccurate estimations and causes the instability of the Lagrange multiplier learning. In this paper, we present a novel off-policy reinforcement learning algorithm called Conservative Distributional Maximum a Posteriori Policy Optimization (CDMPO). At first, to accurately judge whether the current situation satisfies the constraints, CDMPO adapts distributional reinforcement learning method to estimate the Q-function and C-function. Then, CDMPO uses a conservative value function loss to reduce the number of violations of constraints during the exploration process. In addition, we utilize Weighted Average Proportional Integral Derivative (WAPID) to update the Lagrange multiplier stably. Empirical results show that the proposed method has fewer violations of constraints in the early exploration process. The final test results also illustrate that our method has better risk control.
format Preprint
id arxiv_https___arxiv_org_abs_2201_07286
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Conservative Distributional Reinforcement Learning with Safety Constraints
Zhang, Hengrui
Lin, Youfang
Han, Sheng
Wang, Shuo
Lv, Kai
Machine Learning
Artificial Intelligence
Safety exploration can be regarded as a constrained Markov decision problem where the expected long-term cost is constrained. Previous off-policy algorithms convert the constrained optimization problem into the corresponding unconstrained dual problem by introducing the Lagrangian relaxation technique. However, the cost function of the above algorithms provides inaccurate estimations and causes the instability of the Lagrange multiplier learning. In this paper, we present a novel off-policy reinforcement learning algorithm called Conservative Distributional Maximum a Posteriori Policy Optimization (CDMPO). At first, to accurately judge whether the current situation satisfies the constraints, CDMPO adapts distributional reinforcement learning method to estimate the Q-function and C-function. Then, CDMPO uses a conservative value function loss to reduce the number of violations of constraints during the exploration process. In addition, we utilize Weighted Average Proportional Integral Derivative (WAPID) to update the Lagrange multiplier stably. Empirical results show that the proposed method has fewer violations of constraints in the early exploration process. The final test results also illustrate that our method has better risk control.
title Conservative Distributional Reinforcement Learning with Safety Constraints
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2201.07286