A Large Deviations Perspective on Policy Gradient Algorithms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jongeneel, Wouter, Kuhn, Daniel, Li, Mengmeng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913374704500736
author Jongeneel, Wouter
Kuhn, Daniel
Li, Mengmeng
author_facet Jongeneel, Wouter
Kuhn, Daniel
Li, Mengmeng
contents Motivated by policy gradient methods in the context of reinforcement learning, we identify a large deviation rate function for the iterates generated by stochastic gradient descent for possibly non-convex objectives satisfying a Polyak-Łojasiewicz condition. Leveraging the contraction principle from large deviations theory, we illustrate the potential of this result by showing how convergence properties of policy gradient with a softmax parametrization and an entropy regularized objective can be naturally extended to a wide spectrum of other policy parametrizations.
format Preprint
id arxiv_https___arxiv_org_abs_2311_07411
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Large Deviations Perspective on Policy Gradient Algorithms
Jongeneel, Wouter
Kuhn, Daniel
Li, Mengmeng
Optimization and Control
Machine Learning
60F10, 90C26
Motivated by policy gradient methods in the context of reinforcement learning, we identify a large deviation rate function for the iterates generated by stochastic gradient descent for possibly non-convex objectives satisfying a Polyak-Łojasiewicz condition. Leveraging the contraction principle from large deviations theory, we illustrate the potential of this result by showing how convergence properties of policy gradient with a softmax parametrization and an entropy regularized objective can be naturally extended to a wide spectrum of other policy parametrizations.
title A Large Deviations Perspective on Policy Gradient Algorithms
topic Optimization and Control
Machine Learning
60F10, 90C26
url https://arxiv.org/abs/2311.07411