Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Montenegro, Alessandro, Mussi, Marco, Papini, Matteo, Metelli, Alberto Maria
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929588404224000
author Montenegro, Alessandro
Mussi, Marco
Papini, Matteo
Metelli, Alberto Maria
author_facet Montenegro, Alessandro
Mussi, Marco
Papini, Matteo
Metelli, Alberto Maria
contents Constrained Reinforcement Learning (CRL) tackles sequential decision-making problems where agents are required to achieve goals by maximizing the expected return while meeting domain-specific constraints, which are often formulated as expected costs. In this setting, policy-based methods are widely used since they come with several advantages when dealing with continuous-control problems. These methods search in the policy space with an action-based or parameter-based exploration strategy, depending on whether they learn directly the parameters of a stochastic policy or those of a stochastic hyperpolicy. In this paper, we propose a general framework for addressing CRL problems via gradient-based primal-dual algorithms, relying on an alternate ascent/descent scheme with dual-variable regularization. We introduce an exploration-agnostic algorithm, called C-PG, which exhibits global last-iterate convergence guarantees under (weak) gradient domination assumptions, improving and generalizing existing results. Then, we design C-PGAE and C-PGPE, the action-based and the parameter-based versions of C-PG, respectively, and we illustrate how they naturally extend to constraints defined in terms of risk measures over the costs, as it is often requested in safety-critical scenarios. Finally, we numerically validate our algorithms on constrained control problems, and compare them with state-of-the-art baselines, demonstrating their effectiveness.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10775
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning
Montenegro, Alessandro
Mussi, Marco
Papini, Matteo
Metelli, Alberto Maria
Machine Learning
Constrained Reinforcement Learning (CRL) tackles sequential decision-making problems where agents are required to achieve goals by maximizing the expected return while meeting domain-specific constraints, which are often formulated as expected costs. In this setting, policy-based methods are widely used since they come with several advantages when dealing with continuous-control problems. These methods search in the policy space with an action-based or parameter-based exploration strategy, depending on whether they learn directly the parameters of a stochastic policy or those of a stochastic hyperpolicy. In this paper, we propose a general framework for addressing CRL problems via gradient-based primal-dual algorithms, relying on an alternate ascent/descent scheme with dual-variable regularization. We introduce an exploration-agnostic algorithm, called C-PG, which exhibits global last-iterate convergence guarantees under (weak) gradient domination assumptions, improving and generalizing existing results. Then, we design C-PGAE and C-PGPE, the action-based and the parameter-based versions of C-PG, respectively, and we illustrate how they naturally extend to constraints defined in terms of risk measures over the costs, as it is often requested in safety-critical scenarios. Finally, we numerically validate our algorithms on constrained control problems, and compare them with state-of-the-art baselines, demonstrating their effectiveness.
title Last-Iterate Global Convergence of Policy Gradients for Constrained Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2407.10775