Unifying Hamilton-Jacobi Reachability and Reinforcement Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Solanki, Prashant, El-Hajj, Isabelle, van Beers, Jasper, van Kampen, Erik-Jan, de Visser, Coen
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918491745943552
author Solanki, Prashant
El-Hajj, Isabelle
van Beers, Jasper
van Kampen, Erik-Jan
de Visser, Coen
author_facet Solanki, Prashant
El-Hajj, Isabelle
van Beers, Jasper
van Kampen, Erik-Jan
de Visser, Coen
contents We unify Hamilton-Jacobi (HJ) reachability and Reinforcement Learning (RL) through a proposed running cost formulation. We prove that the resultant travel-cost value function is the unique bounded viscosity solution of a time-dependent Hamilton-Jacobi Bellman (HJB) Partial Differential Equation (PDE) with zero terminal data, whose negative sublevel set equals the strict backward-reachable tube. Using a forward reparameterization and a contraction inducing Bellman update, we show that fixed points of small-step RL value iteration converge to the viscosity solution of the forward discounted HJB. Experiments on a classical benchmark validate this connection by demonstrating convergence of learned value functions toward semi-Lagrangian HJB solutions and by quantifying approximation error across the state space. These results empirically support the theoretical analysis, showing that the proposed framework preserves reachability-based safety semantics while remaining compatible with deep RL implementations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08050
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unifying Hamilton-Jacobi Reachability and Reinforcement Learning
Solanki, Prashant
El-Hajj, Isabelle
van Beers, Jasper
van Kampen, Erik-Jan
de Visser, Coen
Systems and Control
We unify Hamilton-Jacobi (HJ) reachability and Reinforcement Learning (RL) through a proposed running cost formulation. We prove that the resultant travel-cost value function is the unique bounded viscosity solution of a time-dependent Hamilton-Jacobi Bellman (HJB) Partial Differential Equation (PDE) with zero terminal data, whose negative sublevel set equals the strict backward-reachable tube. Using a forward reparameterization and a contraction inducing Bellman update, we show that fixed points of small-step RL value iteration converge to the viscosity solution of the forward discounted HJB. Experiments on a classical benchmark validate this connection by demonstrating convergence of learned value functions toward semi-Lagrangian HJB solutions and by quantifying approximation error across the state space. These results empirically support the theoretical analysis, showing that the proposed framework preserves reachability-based safety semantics while remaining compatible with deep RL implementations.
title Unifying Hamilton-Jacobi Reachability and Reinforcement Learning
topic Systems and Control
url https://arxiv.org/abs/2601.08050