Policy Gradient Converges to the Globally Optimal Policy for Nearly Linear-Quadratic Regulators

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Han, Yinbin, Razaviyayn, Meisam, Xu, Renyuan
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908310085566464
author Han, Yinbin
Razaviyayn, Meisam
Xu, Renyuan
author_facet Han, Yinbin
Razaviyayn, Meisam
Xu, Renyuan
contents Nonlinear control systems with partial information to the decision maker are prevalent in a variety of applications. As a step toward studying such nonlinear systems, this work explores reinforcement learning methods for finding the optimal policy in the nearly linear-quadratic regulator systems. In particular, we consider a dynamic system that combines linear and nonlinear components, and is governed by a policy with the same structure. Assuming that the nonlinear component comprises kernels with small Lipschitz coefficients, we characterize the optimization landscape of the cost function. Although the cost function is nonconvex in general, we establish the local strong convexity and smoothness in the vicinity of the global optimizer. Additionally, we propose an initialization mechanism to leverage these properties. Building on the developments, we design a policy gradient algorithm that is guaranteed to converge to the globally optimal policy with a linear rate.
format Preprint
id arxiv_https___arxiv_org_abs_2303_08431
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Policy Gradient Converges to the Globally Optimal Policy for Nearly Linear-Quadratic Regulators
Han, Yinbin
Razaviyayn, Meisam
Xu, Renyuan
Machine Learning
Optimization and Control
Nonlinear control systems with partial information to the decision maker are prevalent in a variety of applications. As a step toward studying such nonlinear systems, this work explores reinforcement learning methods for finding the optimal policy in the nearly linear-quadratic regulator systems. In particular, we consider a dynamic system that combines linear and nonlinear components, and is governed by a policy with the same structure. Assuming that the nonlinear component comprises kernels with small Lipschitz coefficients, we characterize the optimization landscape of the cost function. Although the cost function is nonconvex in general, we establish the local strong convexity and smoothness in the vicinity of the global optimizer. Additionally, we propose an initialization mechanism to leverage these properties. Building on the developments, we design a policy gradient algorithm that is guaranteed to converge to the globally optimal policy with a linear rate.
title Policy Gradient Converges to the Globally Optimal Policy for Nearly Linear-Quadratic Regulators
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2303.08431