Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2504.16417 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912711537852416 |
|---|---|
| author | Mestres, Pol Marzabal, Arnau Cortés, Jorge |
| author_facet | Mestres, Pol Marzabal, Arnau Cortés, Jorge |
| contents | This paper considers the problem of solving constrained
reinforcement learning problems with anytime guarantees, meaning
that the algorithmic solution returns a safe policy regardless of
when it is terminated. Drawing inspiration from anytime constrained
optimization, we introduce Reinforcement Learning-based Safe
Gradient Flow (RL-SGF), an on-policy algorithm which employs
estimates of the value functions and their respective gradients
associated with the objective and safety constraints for the current
policy, and updates the policy parameters by solving a convex
quadratically constrained quadratic program. We show that if the
estimates are computed with a sufficiently large number of episodes
(for which we provide an explicit bound), safe policies are updated
to safe policies with a probability higher than a prescribed
tolerance. We also show that iterates asymptotically converge to a
neighborhood of a KKT point, whose size can be arbitrarily reduced
by refining the estimates of the value function and their gradients.
We illustrate the performance of RL-SGF in a navigation example. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_16417 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Anytime Safe Reinforcement Learning Mestres, Pol Marzabal, Arnau Cortés, Jorge Systems and Control This paper considers the problem of solving constrained reinforcement learning problems with anytime guarantees, meaning that the algorithmic solution returns a safe policy regardless of when it is terminated. Drawing inspiration from anytime constrained optimization, we introduce Reinforcement Learning-based Safe Gradient Flow (RL-SGF), an on-policy algorithm which employs estimates of the value functions and their respective gradients associated with the objective and safety constraints for the current policy, and updates the policy parameters by solving a convex quadratically constrained quadratic program. We show that if the estimates are computed with a sufficiently large number of episodes (for which we provide an explicit bound), safe policies are updated to safe policies with a probability higher than a prescribed tolerance. We also show that iterates asymptotically converge to a neighborhood of a KKT point, whose size can be arbitrarily reduced by refining the estimates of the value function and their gradients. We illustrate the performance of RL-SGF in a navigation example. |
| title | Anytime Safe Reinforcement Learning |
| topic | Systems and Control |
| url | https://arxiv.org/abs/2504.16417 |