A universal policy wrapper with guarantees

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bolychev, Anton, Malaniya, Georgiy, Yaremenko, Grigory, Krasnaya, Anastasia, Osinenko, Pavel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910951487307776
author Bolychev, Anton
Malaniya, Georgiy
Yaremenko, Grigory
Krasnaya, Anastasia
Osinenko, Pavel
author_facet Bolychev, Anton
Malaniya, Georgiy
Yaremenko, Grigory
Krasnaya, Anastasia
Osinenko, Pavel
contents We introduce a universal policy wrapper for reinforcement learning agents that ensures formal goal-reaching guarantees. In contrast to standard reinforcement learning algorithms that excel in performance but lack rigorous safety assurances, our wrapper selectively switches between a high-performing base policy -- derived from any existing RL method -- and a fallback policy with known convergence properties. Base policy's value function supervises this switching process, determining when the fallback policy should override the base policy to ensure the system remains on a stable path. The analysis proves that our wrapper inherits the fallback policy's goal-reaching guarantees while preserving or improving upon the performance of the base policy. Notably, it operates without needing additional system knowledge or online constrained optimization, making it readily deployable across diverse reinforcement learning architectures and tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12354
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A universal policy wrapper with guarantees
Bolychev, Anton
Malaniya, Georgiy
Yaremenko, Grigory
Krasnaya, Anastasia
Osinenko, Pavel
Machine Learning
Artificial Intelligence
Robotics
Systems and Control
Optimization and Control
We introduce a universal policy wrapper for reinforcement learning agents that ensures formal goal-reaching guarantees. In contrast to standard reinforcement learning algorithms that excel in performance but lack rigorous safety assurances, our wrapper selectively switches between a high-performing base policy -- derived from any existing RL method -- and a fallback policy with known convergence properties. Base policy's value function supervises this switching process, determining when the fallback policy should override the base policy to ensure the system remains on a stable path. The analysis proves that our wrapper inherits the fallback policy's goal-reaching guarantees while preserving or improving upon the performance of the base policy. Notably, it operates without needing additional system knowledge or online constrained optimization, making it readily deployable across diverse reinforcement learning architectures and tasks.
title A universal policy wrapper with guarantees
topic Machine Learning
Artificial Intelligence
Robotics
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2505.12354