Convex Regularization and Convergence of Policy Gradient Flows under Safety Constraints

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Malo, Pekka, Viitasaari, Lauri, Suominen, Antti, Vilkkumaa, Eeva, Tahvonen, Olli
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916951293427712
author Malo, Pekka
Viitasaari, Lauri
Suominen, Antti
Vilkkumaa, Eeva
Tahvonen, Olli
author_facet Malo, Pekka
Viitasaari, Lauri
Suominen, Antti
Vilkkumaa, Eeva
Tahvonen, Olli
contents This paper examines reinforcement learning (RL) in infinite-horizon decision processes with almost-sure safety constraints, crucial for applications like autonomous systems, finance, and resource management. We propose a doubly-regularized RL framework combining reward and parameter regularization to address safety constraints in continuous state-action spaces. The problem is formulated as a convex regularized objective with parametrized policies in the mean-field regime. Leveraging mean-field theory and Wasserstein gradient flows, policies are modeled on an infinite-dimensional statistical manifold, with updates governed by parameter distribution gradient flows. Key contributions include solvability conditions for safety-constrained problems, smooth bounded approximations for gradient flows, and exponential convergence guarantees under sufficient regularization. General regularization conditions, including entropy regularization, support practical particle method implementations. This framework provides robust theoretical insights and guarantees for safe RL in complex, high-dimensional settings.
format Preprint
id arxiv_https___arxiv_org_abs_2411_19193
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Convex Regularization and Convergence of Policy Gradient Flows under Safety Constraints
Malo, Pekka
Viitasaari, Lauri
Suominen, Antti
Vilkkumaa, Eeva
Tahvonen, Olli
Machine Learning
Artificial Intelligence
Optimization and Control
Probability
90C26, 90C40, 90C46, 93E20, 60B05
This paper examines reinforcement learning (RL) in infinite-horizon decision processes with almost-sure safety constraints, crucial for applications like autonomous systems, finance, and resource management. We propose a doubly-regularized RL framework combining reward and parameter regularization to address safety constraints in continuous state-action spaces. The problem is formulated as a convex regularized objective with parametrized policies in the mean-field regime. Leveraging mean-field theory and Wasserstein gradient flows, policies are modeled on an infinite-dimensional statistical manifold, with updates governed by parameter distribution gradient flows. Key contributions include solvability conditions for safety-constrained problems, smooth bounded approximations for gradient flows, and exponential convergence guarantees under sufficient regularization. General regularization conditions, including entropy regularization, support practical particle method implementations. This framework provides robust theoretical insights and guarantees for safe RL in complex, high-dimensional settings.
title Convex Regularization and Convergence of Policy Gradient Flows under Safety Constraints
topic Machine Learning
Artificial Intelligence
Optimization and Control
Probability
90C26, 90C40, 90C46, 93E20, 60B05
url https://arxiv.org/abs/2411.19193