Learning to Provably Satisfy High Relative Degree Constraints for Black-Box Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bouvier, Jean-Baptiste, Nagpal, Kartik, Mehr, Negar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910546085806080
author Bouvier, Jean-Baptiste
Nagpal, Kartik
Mehr, Negar
author_facet Bouvier, Jean-Baptiste
Nagpal, Kartik
Mehr, Negar
contents In this paper, we develop a method for learning a control policy guaranteed to satisfy an affine state constraint of high relative degree in closed loop with a black-box system. Previous reinforcement learning (RL) approaches to satisfy safety constraints either require access to the system model, or assume control affine dynamics, or only discourage violations with reward shaping. Only recently have these issues been addressed with POLICEd RL, which guarantees constraint satisfaction for black-box systems. However, this previous work can only enforce constraints of relative degree 1. To address this gap, we build a novel RL algorithm explicitly designed to enforce an affine state constraint of high relative degree in closed loop with a black-box control system. Our key insight is to make the learned policy be affine around the unsafe set and to use this affine region to dissipate the inertia of the high relative degree constraint. We prove that such policies guarantee constraint satisfaction for deterministic systems while being agnostic to the choice of the RL training algorithm. Our results demonstrate the capacity of our approach to enforce hard constraints in the Gym inverted pendulum and on a space shuttle landing simulation.
format Preprint
id arxiv_https___arxiv_org_abs_2407_20456
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning to Provably Satisfy High Relative Degree Constraints for Black-Box Systems
Bouvier, Jean-Baptiste
Nagpal, Kartik
Mehr, Negar
Systems and Control
In this paper, we develop a method for learning a control policy guaranteed to satisfy an affine state constraint of high relative degree in closed loop with a black-box system. Previous reinforcement learning (RL) approaches to satisfy safety constraints either require access to the system model, or assume control affine dynamics, or only discourage violations with reward shaping. Only recently have these issues been addressed with POLICEd RL, which guarantees constraint satisfaction for black-box systems. However, this previous work can only enforce constraints of relative degree 1. To address this gap, we build a novel RL algorithm explicitly designed to enforce an affine state constraint of high relative degree in closed loop with a black-box control system. Our key insight is to make the learned policy be affine around the unsafe set and to use this affine region to dissipate the inertia of the high relative degree constraint. We prove that such policies guarantee constraint satisfaction for deterministic systems while being agnostic to the choice of the RL training algorithm. Our results demonstrate the capacity of our approach to enforce hard constraints in the Gym inverted pendulum and on a space shuttle landing simulation.
title Learning to Provably Satisfy High Relative Degree Constraints for Black-Box Systems
topic Systems and Control
url https://arxiv.org/abs/2407.20456