A Provable Approach for End-to-End Safe Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wachi, Akifumi, Miyaguchi, Kohei, Tanabe, Takumi, Sato, Rei, Akimoto, Youhei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916762956595200
author Wachi, Akifumi
Miyaguchi, Kohei
Tanabe, Takumi
Sato, Rei
Akimoto, Youhei
author_facet Wachi, Akifumi
Miyaguchi, Kohei
Tanabe, Takumi
Sato, Rei
Akimoto, Youhei
contents A longstanding goal in safe reinforcement learning (RL) is a method to ensure the safety of a policy throughout the entire process, from learning to operation. However, existing safe RL paradigms inherently struggle to achieve this objective. We propose a method, called Provably Lifetime Safe RL (PLS), that integrates offline safe RL with safe policy deployment to address this challenge. Our proposed method learns a policy offline using return-conditioned supervised learning and then deploys the resulting policy while cautiously optimizing a limited set of parameters, known as target returns, using Gaussian processes (GPs). Theoretically, we justify the use of GPs by analyzing the mathematical relationship between target and actual returns. We then prove that PLS finds near-optimal target returns while guaranteeing safety with high probability. Empirically, we demonstrate that PLS outperforms baselines both in safety and reward performance, thereby achieving the longstanding goal to obtain high rewards while ensuring the safety of a policy throughout the lifetime from learning to operation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Provable Approach for End-to-End Safe Reinforcement Learning
Wachi, Akifumi
Miyaguchi, Kohei
Tanabe, Takumi
Sato, Rei
Akimoto, Youhei
Machine Learning
Artificial Intelligence
Information Theory
Robotics
A longstanding goal in safe reinforcement learning (RL) is a method to ensure the safety of a policy throughout the entire process, from learning to operation. However, existing safe RL paradigms inherently struggle to achieve this objective. We propose a method, called Provably Lifetime Safe RL (PLS), that integrates offline safe RL with safe policy deployment to address this challenge. Our proposed method learns a policy offline using return-conditioned supervised learning and then deploys the resulting policy while cautiously optimizing a limited set of parameters, known as target returns, using Gaussian processes (GPs). Theoretically, we justify the use of GPs by analyzing the mathematical relationship between target and actual returns. We then prove that PLS finds near-optimal target returns while guaranteeing safety with high probability. Empirically, we demonstrate that PLS outperforms baselines both in safety and reward performance, thereby achieving the longstanding goal to obtain high rewards while ensuring the safety of a policy throughout the lifetime from learning to operation.
title A Provable Approach for End-to-End Safe Reinforcement Learning
topic Machine Learning
Artificial Intelligence
Information Theory
Robotics
url https://arxiv.org/abs/2505.21852