Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Qisong, Huang, Xinmiao, Hu, Jinwei, Li, Zhuoyun, Dong, Yi, Wu, Changshun, Huang, Xiaowei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913126529630208
author He, Qisong
Huang, Xinmiao
Hu, Jinwei
Li, Zhuoyun
Dong, Yi
Wu, Changshun
Huang, Xiaowei
author_facet He, Qisong
Huang, Xinmiao
Hu, Jinwei
Li, Zhuoyun
Dong, Yi
Wu, Changshun
Huang, Xiaowei
contents Safe navigation for mobile robots demands policies that remain reliable under the high-consequence perception uncertainty of cluttered environments. Yet most existing safe reinforcement learning (RL) methods assess safety through average cumulative cost. Such metrics can mask dangerous tail-risk behaviors. To address this, we propose a framework that trains risk-sensitive policies through Conditional Value-at-Risk (CVaR) constrained optimization on an off-policy TD3 backbone and evaluates their safety margins post-training through neural network reachability verification. During training, the policy is optimized under CVaR constraints on cumulative costs, promoting sensitivity to high-cost tail outcomes rather than average behavior alone. After training, we compute action reachable sets under bounded observation uncertainty using Taylor Model analysis, yielding a safety rate metric that quantifies the proportion of evaluated states at which the policy's reachable action set remains within prescribed safety margins. A key finding is that policies trained with CVaR constraints maintain larger safety margins from obstacles across evaluated states. This makes them significantly more amenable to formal reachability verification. Experiments across ten navigation scenarios and six baselines show that our method achieves a 98.3\% success rate, the highest safety verification rate among all compared methods, while revealing that average cost rankings and reachability-based safety rankings can diverge. This indicates that reachability verification captures risks which are missed by empirical cost metrics alone. We further validate our approach on a physical Clearpath Jackal robot, demonstrating successful sim-to-real transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14174
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation
He, Qisong
Huang, Xinmiao
Hu, Jinwei
Li, Zhuoyun
Dong, Yi
Wu, Changshun
Huang, Xiaowei
Robotics
Safe navigation for mobile robots demands policies that remain reliable under the high-consequence perception uncertainty of cluttered environments. Yet most existing safe reinforcement learning (RL) methods assess safety through average cumulative cost. Such metrics can mask dangerous tail-risk behaviors. To address this, we propose a framework that trains risk-sensitive policies through Conditional Value-at-Risk (CVaR) constrained optimization on an off-policy TD3 backbone and evaluates their safety margins post-training through neural network reachability verification. During training, the policy is optimized under CVaR constraints on cumulative costs, promoting sensitivity to high-cost tail outcomes rather than average behavior alone. After training, we compute action reachable sets under bounded observation uncertainty using Taylor Model analysis, yielding a safety rate metric that quantifies the proportion of evaluated states at which the policy's reachable action set remains within prescribed safety margins. A key finding is that policies trained with CVaR constraints maintain larger safety margins from obstacles across evaluated states. This makes them significantly more amenable to formal reachability verification. Experiments across ten navigation scenarios and six baselines show that our method achieves a 98.3\% success rate, the highest safety verification rate among all compared methods, while revealing that average cost rankings and reachability-based safety rankings can diverge. This indicates that reachability verification captures risks which are missed by empirical cost metrics alone. We further validate our approach on a physical Clearpath Jackal robot, demonstrating successful sim-to-real transfer.
title Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation
topic Robotics
url https://arxiv.org/abs/2605.14174