Neural Policy Iteration for Stochastic Optimal Control: A Physics-Informed Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Yeongjong, Kim, Yeoneung, Kim, Minseok, Cho, Namkyeong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915423641927680
author Kim, Yeongjong
Kim, Yeoneung
Kim, Minseok
Cho, Namkyeong
author_facet Kim, Yeongjong
Kim, Yeoneung
Kim, Minseok
Cho, Namkyeong
contents We propose a physics-informed neural network policy iteration (PINN-PI) framework for solving stochastic optimal control problems governed by second-order Hamilton--Jacobi--Bellman (HJB) equations. At each iteration, a neural network is trained to approximate the value function by minimizing the residual of a linear PDE induced by a fixed policy. This linear structure enables systematic $L^2$ error control at each policy evaluation step, and allows us to derive explicit Lipschitz-type bounds that quantify how value gradient errors propagate to the policy updates. This interpretability provides a theoretical basis for evaluating policy quality during training. Our method extends recent deterministic PINN-based approaches to stochastic settings, inheriting the global exponential convergence guarantees of classical policy iteration under mild conditions. We demonstrate the effectiveness of our method on several benchmark problems, including stochastic cartpole, pendulum problems and high-dimensional linear quadratic regulation (LQR) problems in up to 10D.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01718
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Neural Policy Iteration for Stochastic Optimal Control: A Physics-Informed Approach
Kim, Yeongjong
Kim, Yeoneung
Kim, Minseok
Cho, Namkyeong
Machine Learning
Computational Engineering, Finance, and Science
Numerical Analysis
93E20, 35Q93, 68T07, 65N21
We propose a physics-informed neural network policy iteration (PINN-PI) framework for solving stochastic optimal control problems governed by second-order Hamilton--Jacobi--Bellman (HJB) equations. At each iteration, a neural network is trained to approximate the value function by minimizing the residual of a linear PDE induced by a fixed policy. This linear structure enables systematic $L^2$ error control at each policy evaluation step, and allows us to derive explicit Lipschitz-type bounds that quantify how value gradient errors propagate to the policy updates. This interpretability provides a theoretical basis for evaluating policy quality during training. Our method extends recent deterministic PINN-based approaches to stochastic settings, inheriting the global exponential convergence guarantees of classical policy iteration under mild conditions. We demonstrate the effectiveness of our method on several benchmark problems, including stochastic cartpole, pendulum problems and high-dimensional linear quadratic regulation (LQR) problems in up to 10D.
title Neural Policy Iteration for Stochastic Optimal Control: A Physics-Informed Approach
topic Machine Learning
Computational Engineering, Finance, and Science
Numerical Analysis
93E20, 35Q93, 68T07, 65N21
url https://arxiv.org/abs/2508.01718