Learn With Imagination: Safe Set Guided State-wise Constrained Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Yifan, Li, Feihan, Zhao, Weiye, Chen, Rui, Wei, Tianhao, Liu, Changliu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908389855985664
author Sun, Yifan
Li, Feihan
Zhao, Weiye
Chen, Rui
Wei, Tianhao
Liu, Changliu
author_facet Sun, Yifan
Li, Feihan
Zhao, Weiye
Chen, Rui
Wei, Tianhao
Liu, Changliu
contents Deep reinforcement learning (RL) excels in various control tasks, yet the absence of safety guarantees hampers its real-world applicability. In particular, explorations during learning usually results in safety violations, while the RL agent learns from those mistakes. On the other hand, safe control techniques ensure persistent safety satisfaction but demand strong priors on system dynamics, which is usually hard to obtain in practice. To address these problems, we present Safe Set Guided State-wise Constrained Policy Optimization (S-3PO), a pioneering algorithm generating state-wise safe optimal policies with zero training violations, i.e., learning without mistakes. S-3PO first employs a safety-oriented monitor with black-box dynamics to ensure safe exploration. It then enforces an "imaginary" cost for the RL agent to converge to optimal behaviors within safety constraints. S-3PO outperforms existing methods in high-dimensional robotics tasks, managing state-wise constraints with zero training violation. This innovation marks a significant stride towards real-world safe RL deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2308_13140
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learn With Imagination: Safe Set Guided State-wise Constrained Policy Optimization
Sun, Yifan
Li, Feihan
Zhao, Weiye
Chen, Rui
Wei, Tianhao
Liu, Changliu
Robotics
Deep reinforcement learning (RL) excels in various control tasks, yet the absence of safety guarantees hampers its real-world applicability. In particular, explorations during learning usually results in safety violations, while the RL agent learns from those mistakes. On the other hand, safe control techniques ensure persistent safety satisfaction but demand strong priors on system dynamics, which is usually hard to obtain in practice. To address these problems, we present Safe Set Guided State-wise Constrained Policy Optimization (S-3PO), a pioneering algorithm generating state-wise safe optimal policies with zero training violations, i.e., learning without mistakes. S-3PO first employs a safety-oriented monitor with black-box dynamics to ensure safe exploration. It then enforces an "imaginary" cost for the RL agent to converge to optimal behaviors within safety constraints. S-3PO outperforms existing methods in high-dimensional robotics tasks, managing state-wise constraints with zero training violation. This innovation marks a significant stride towards real-world safe RL deployment.
title Learn With Imagination: Safe Set Guided State-wise Constrained Policy Optimization
topic Robotics
url https://arxiv.org/abs/2308.13140