Entropy annealing for policy mirror descent in continuous time and space

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sethi, Deven, Šiška, David, Zhang, Yufei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910889016295424
author Sethi, Deven
Šiška, David
Zhang, Yufei
author_facet Sethi, Deven
Šiška, David
Zhang, Yufei
contents Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additional regularization bias. This work quantifies the impact of entropy regularization on the convergence of policy gradient methods for stochastic exit time control problems. We analyze a continuous-time policy mirror descent dynamics, which updates the policy based on the gradient of an entropy-regularized value function and adjusts the strength of entropy regularization as the algorithm progresses. We prove that with a fixed entropy level, the mirror descent dynamics converges exponentially to the optimal solution of the regularized problem. We further show that when the entropy level decays at suitable polynomial rates, the annealed flow converges to the solution of the unregularized problem at a rate of $\mathcal O(1/S)$ for discrete action spaces and, under suitable conditions, at a rate of $\mathcal O(1/\sqrt{S})$ for general action spaces, with $S$ being the gradient flow running time. The technical challenge lies in analyzing the gradient flow in the infinite-dimensional space of Markov kernels for nonconvex objectives. This paper explains how entropy regularization improves policy optimization, even with the true gradient, from the perspective of convergence rate.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20250
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Entropy annealing for policy mirror descent in continuous time and space
Sethi, Deven
Šiška, David
Zhang, Yufei
Optimization and Control
Machine Learning
Probability
Primary 93E20, Secondary 49M29, 68Q25, 60H30, 35J61
Entropy regularization has been widely used in policy optimization algorithms to enhance exploration and the robustness of the optimal control; however it also introduces an additional regularization bias. This work quantifies the impact of entropy regularization on the convergence of policy gradient methods for stochastic exit time control problems. We analyze a continuous-time policy mirror descent dynamics, which updates the policy based on the gradient of an entropy-regularized value function and adjusts the strength of entropy regularization as the algorithm progresses. We prove that with a fixed entropy level, the mirror descent dynamics converges exponentially to the optimal solution of the regularized problem. We further show that when the entropy level decays at suitable polynomial rates, the annealed flow converges to the solution of the unregularized problem at a rate of $\mathcal O(1/S)$ for discrete action spaces and, under suitable conditions, at a rate of $\mathcal O(1/\sqrt{S})$ for general action spaces, with $S$ being the gradient flow running time. The technical challenge lies in analyzing the gradient flow in the infinite-dimensional space of Markov kernels for nonconvex objectives. This paper explains how entropy regularization improves policy optimization, even with the true gradient, from the perspective of convergence rate.
title Entropy annealing for policy mirror descent in continuous time and space
topic Optimization and Control
Machine Learning
Probability
Primary 93E20, Secondary 49M29, 68Q25, 60H30, 35J61
url https://arxiv.org/abs/2405.20250