Do No Harm: A Counterfactual Approach to Safe Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vaskov, Sean, Schwarting, Wilko, Baker, Chris L.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911881969532928
author Vaskov, Sean
Schwarting, Wilko
Baker, Chris L.
author_facet Vaskov, Sean
Schwarting, Wilko
Baker, Chris L.
contents Reinforcement Learning (RL) for control has become increasingly popular due to its ability to learn rich feedback policies that take into account uncertainty and complex representations of the environment. When considering safety constraints, constrained optimization approaches, where agents are penalized for constraint violations, are commonly used. In such methods, if agents are initialized in, or must visit, states where constraint violation might be inevitable, it is unclear how much they should be penalized. We address this challenge by formulating a constraint on the counterfactual harm of the learned policy compared to a default, safe policy. In a philosophical sense this formulation only penalizes the learner for constraint violations that it caused; in a practical sense it maintains feasibility of the optimal control problem. We present simulation studies on a rover with uncertain road friction and a tractor-trailer parking environment that demonstrate our constraint formulation enables agents to learn safer policies than contemporary constrained RL methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11669
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do No Harm: A Counterfactual Approach to Safe Reinforcement Learning
Vaskov, Sean
Schwarting, Wilko
Baker, Chris L.
Machine Learning
Artificial Intelligence
Reinforcement Learning (RL) for control has become increasingly popular due to its ability to learn rich feedback policies that take into account uncertainty and complex representations of the environment. When considering safety constraints, constrained optimization approaches, where agents are penalized for constraint violations, are commonly used. In such methods, if agents are initialized in, or must visit, states where constraint violation might be inevitable, it is unclear how much they should be penalized. We address this challenge by formulating a constraint on the counterfactual harm of the learned policy compared to a default, safe policy. In a philosophical sense this formulation only penalizes the learner for constraint violations that it caused; in a practical sense it maintains feasibility of the optimal control problem. We present simulation studies on a rover with uncertain road friction and a tractor-trailer parking environment that demonstrate our constraint formulation enables agents to learn safer policies than contemporary constrained RL methods.
title Do No Harm: A Counterfactual Approach to Safe Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.11669