Convergence of Policy Mirror Descent Beyond Compatible Function Approximation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sherman, Uri, Koren, Tomer, Mansour, Yishay
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913929074049024
author Sherman, Uri
Koren, Tomer
Mansour, Yishay
author_facet Sherman, Uri
Koren, Tomer
Mansour, Yishay
contents Modern policy optimization methods roughly follow the policy mirror descent (PMD) algorithmic template, for which there are by now numerous theoretical convergence results. However, most of these either target tabular environments, or can be applied effectively only when the class of policies being optimized over satisfies strong closure conditions, which is typically not the case when working with parametric policy classes in large-scale environments. In this work, we develop a theoretical framework for PMD for general policy classes where we replace the closure conditions with a strictly weaker variational gradient dominance assumption, and obtain upper bounds on the rate of convergence to the best-in-class policy. Our main result leverages a novel notion of smoothness with respect to a local norm induced by the occupancy measure of the current policy, and casts PMD as a particular instance of smooth non-convex optimization in non-Euclidean space.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11033
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
Sherman, Uri
Koren, Tomer
Mansour, Yishay
Machine Learning
Optimization and Control
Modern policy optimization methods roughly follow the policy mirror descent (PMD) algorithmic template, for which there are by now numerous theoretical convergence results. However, most of these either target tabular environments, or can be applied effectively only when the class of policies being optimized over satisfies strong closure conditions, which is typically not the case when working with parametric policy classes in large-scale environments. In this work, we develop a theoretical framework for PMD for general policy classes where we replace the closure conditions with a strictly weaker variational gradient dominance assumption, and obtain upper bounds on the rate of convergence to the best-in-class policy. Our main result leverages a novel notion of smoothness with respect to a local norm induced by the occupancy measure of the current policy, and casts PMD as a particular instance of smooth non-convex optimization in non-Euclidean space.
title Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2502.11033