Stacked Universal Successor Feature Approximators for Safety in Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cannon, Ian, Garcia, Washington, Gresavage, Thomas, Saurine, Joseph, Leong, Ian, Culbertson, Jared
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909307918876672
author Cannon, Ian
Garcia, Washington
Gresavage, Thomas
Saurine, Joseph
Leong, Ian
Culbertson, Jared
author_facet Cannon, Ian
Garcia, Washington
Gresavage, Thomas
Saurine, Joseph
Leong, Ian
Culbertson, Jared
contents Real-world problems often involve complex objective structures that resist distillation into reinforcement learning environments with a single objective. Operation costs must be balanced with multi-dimensional task performance and end-states' effects on future availability, all while ensuring safety for other agents in the environment and the reinforcement learning agent itself. System redundancy through secondary backup controllers has proven to be an effective method to ensure safety in real-world applications where the risk of violating constraints is extremely high. In this work, we investigate the utility of a stacked, continuous-control variation of universal successor feature approximation (USFA) adapted for soft actor-critic (SAC) and coupled with a suite of secondary safety controllers, which we call stacked USFA for safety (SUSFAS). Our method improves performance on secondary objectives compared to SAC baselines using an intervening secondary controller such as a runtime assurance (RTA) controller.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04641
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stacked Universal Successor Feature Approximators for Safety in Reinforcement Learning
Cannon, Ian
Garcia, Washington
Gresavage, Thomas
Saurine, Joseph
Leong, Ian
Culbertson, Jared
Machine Learning
Artificial Intelligence
Real-world problems often involve complex objective structures that resist distillation into reinforcement learning environments with a single objective. Operation costs must be balanced with multi-dimensional task performance and end-states' effects on future availability, all while ensuring safety for other agents in the environment and the reinforcement learning agent itself. System redundancy through secondary backup controllers has proven to be an effective method to ensure safety in real-world applications where the risk of violating constraints is extremely high. In this work, we investigate the utility of a stacked, continuous-control variation of universal successor feature approximation (USFA) adapted for soft actor-critic (SAC) and coupled with a suite of secondary safety controllers, which we call stacked USFA for safety (SUSFAS). Our method improves performance on secondary objectives compared to SAC baselines using an intervening secondary controller such as a runtime assurance (RTA) controller.
title Stacked Universal Successor Feature Approximators for Safety in Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2409.04641