A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alfano, Carlo, Yuan, Rui, Rebeschini, Patrick
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917588197441536
author Alfano, Carlo
Yuan, Rui
Rebeschini, Patrick
author_facet Alfano, Carlo
Yuan, Rui
Rebeschini, Patrick
contents Modern policy optimization methods in reinforcement learning, such as TRPO and PPO, owe their success to the use of parameterized policies. However, while theoretical guarantees have been established for this class of algorithms, especially in the tabular setting, the use of general parameterization schemes remains mostly unjustified. In this work, we introduce a novel framework for policy optimization based on mirror descent that naturally accommodates general parameterizations. The policy class induced by our scheme recovers known classes, e.g., softmax, and generates new ones depending on the choice of mirror map. Using our framework, we obtain the first result that guarantees linear convergence for a policy-gradient-based method involving general parameterization. To demonstrate the ability of our framework to accommodate general parameterization schemes, we provide its sample complexity when using shallow neural networks, show that it represents an improvement upon the previous best results, and empirically validate the effectiveness of our theoretical claims on classic control tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2301_13139
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence
Alfano, Carlo
Yuan, Rui
Rebeschini, Patrick
Machine Learning
Optimization and Control
Statistics Theory
Modern policy optimization methods in reinforcement learning, such as TRPO and PPO, owe their success to the use of parameterized policies. However, while theoretical guarantees have been established for this class of algorithms, especially in the tabular setting, the use of general parameterization schemes remains mostly unjustified. In this work, we introduce a novel framework for policy optimization based on mirror descent that naturally accommodates general parameterizations. The policy class induced by our scheme recovers known classes, e.g., softmax, and generates new ones depending on the choice of mirror map. Using our framework, we obtain the first result that guarantees linear convergence for a policy-gradient-based method involving general parameterization. To demonstrate the ability of our framework to accommodate general parameterization schemes, we provide its sample complexity when using shallow neural networks, show that it represents an improvement upon the previous best results, and empirically validate the effectiveness of our theoretical claims on classic control tasks.
title A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence
topic Machine Learning
Optimization and Control
Statistics Theory
url https://arxiv.org/abs/2301.13139