Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yao, Yihang, Liu, Zuxin, Cen, Zhepeng, Zhu, Jiacheng, Yu, Wenhao, Zhang, Tingnan, Zhao, Ding
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914775967989760
author Yao, Yihang
Liu, Zuxin
Cen, Zhepeng
Zhu, Jiacheng
Yu, Wenhao
Zhang, Tingnan
Zhao, Ding
author_facet Yao, Yihang
Liu, Zuxin
Cen, Zhepeng
Zhu, Jiacheng
Yu, Wenhao
Zhang, Tingnan
Zhao, Ding
contents Safe reinforcement learning (RL) focuses on training reward-maximizing agents subject to pre-defined safety constraints. Yet, learning versatile safe policies that can adapt to varying safety constraint requirements during deployment without retraining remains a largely unexplored and challenging area. In this work, we formulate the versatile safe RL problem and consider two primary requirements: training efficiency and zero-shot adaptation capability. To address them, we introduce the Conditioned Constrained Policy Optimization (CCPO) framework, consisting of two key modules: (1) Versatile Value Estimation (VVE) for approximating value functions under unseen threshold conditions, and (2) Conditioned Variational Inference (CVI) for encoding arbitrary constraint thresholds during policy optimization. Our extensive experiments demonstrate that CCPO outperforms the baselines in terms of safety and task performance while preserving zero-shot adaptation capabilities to different constraint thresholds data-efficiently. This makes our approach suitable for real-world dynamic applications.
format Preprint
id arxiv_https___arxiv_org_abs_2310_03718
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
Yao, Yihang
Liu, Zuxin
Cen, Zhepeng
Zhu, Jiacheng
Yu, Wenhao
Zhang, Tingnan
Zhao, Ding
Machine Learning
Artificial Intelligence
Safe reinforcement learning (RL) focuses on training reward-maximizing agents subject to pre-defined safety constraints. Yet, learning versatile safe policies that can adapt to varying safety constraint requirements during deployment without retraining remains a largely unexplored and challenging area. In this work, we formulate the versatile safe RL problem and consider two primary requirements: training efficiency and zero-shot adaptation capability. To address them, we introduce the Conditioned Constrained Policy Optimization (CCPO) framework, consisting of two key modules: (1) Versatile Value Estimation (VVE) for approximating value functions under unseen threshold conditions, and (2) Conditioned Variational Inference (CVI) for encoding arbitrary constraint thresholds during policy optimization. Our extensive experiments demonstrate that CCPO outperforms the baselines in terms of safety and task performance while preserving zero-shot adaptation capabilities to different constraint thresholds data-efficiently. This makes our approach suitable for real-world dynamic applications.
title Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2310.03718