On the Convergence of the Policy Iteration for Infinite-Horizon Nonlinear Optimal Control Problems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ehring, Tobias, Azmi, Behzad, Haasdonk, Bernard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912481252737024
author Ehring, Tobias
Azmi, Behzad
Haasdonk, Bernard
author_facet Ehring, Tobias
Azmi, Behzad
Haasdonk, Bernard
contents Policy iteration (PI) is a widely used algorithm for synthesizing optimal feedback control policies across many engineering and scientific applications. When PI is deployed on infinite-horizon, nonlinear, autonomous optimal-control problems, however, a number of significant theoretical challenges emerge - particularly when the computational state space is restricted to a bounded domain. In this paper, we investigate these challenges and show that the viability of PI in this setting hinges on the existence, uniqueness, and regularity of solutions to the Generalized Hamilton-Jacobi-Bellman (GHJB) equation solved at each iteration. To ensure a well-posed iterative scheme, the GHJB solution must possess sufficient smoothness, and the domain on which the GHJB equation is solved must remain forward-invariant under the closed-loop dynamics induced by the current policy. Although fundamental to the method's convergence, previous studies have largely overlooked these aspects. This paper closes that gap by introducing a constructive procedure that guarantees forward invariance of the computational domain throughout the entire PI sequence and by establishing sufficient conditions under which a suitably regular GHJB solution exists at every iteration. Numerical results are presented for a grid-based implementation of PI to support the theoretical findings.
format Preprint
id arxiv_https___arxiv_org_abs_2507_09994
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the Convergence of the Policy Iteration for Infinite-Horizon Nonlinear Optimal Control Problems
Ehring, Tobias
Azmi, Behzad
Haasdonk, Bernard
Optimization and Control
Numerical Analysis
49L20, 49N35, 49J15, 49L12
Policy iteration (PI) is a widely used algorithm for synthesizing optimal feedback control policies across many engineering and scientific applications. When PI is deployed on infinite-horizon, nonlinear, autonomous optimal-control problems, however, a number of significant theoretical challenges emerge - particularly when the computational state space is restricted to a bounded domain. In this paper, we investigate these challenges and show that the viability of PI in this setting hinges on the existence, uniqueness, and regularity of solutions to the Generalized Hamilton-Jacobi-Bellman (GHJB) equation solved at each iteration. To ensure a well-posed iterative scheme, the GHJB solution must possess sufficient smoothness, and the domain on which the GHJB equation is solved must remain forward-invariant under the closed-loop dynamics induced by the current policy. Although fundamental to the method's convergence, previous studies have largely overlooked these aspects. This paper closes that gap by introducing a constructive procedure that guarantees forward invariance of the computational domain throughout the entire PI sequence and by establishing sufficient conditions under which a suitably regular GHJB solution exists at every iteration. Numerical results are presented for a grid-based implementation of PI to support the theoretical findings.
title On the Convergence of the Policy Iteration for Infinite-Horizon Nonlinear Optimal Control Problems
topic Optimization and Control
Numerical Analysis
49L20, 49N35, 49J15, 49L12
url https://arxiv.org/abs/2507.09994