Learning-based primal-dual optimal control of discrete-time stochastic systems with multiplicative noise

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Xiushan, Zhang, Weihai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915320944394240
author Jiang, Xiushan
Zhang, Weihai
author_facet Jiang, Xiushan
Zhang, Weihai
contents Reinforcement learning (RL) is an effective approach for solving optimal control problems without knowing the exact information of the system model. However, the classical Q-learning method, a model-free RL algorithm, has its limitations, such as lack of strict theoretical analysis and the need for artificial disturbances during implementation. This paper explores the partially model-free stochastic linear quadratic regulator (SLQR) problem for a system with multiplicative noise from the primal-dual perspective to address these challenges. This approach lays a strong theoretical foundation for understanding the intrinsic mechanisms of classical RL algorithms. We reformulate the SLQR into a non-convex primal-dual optimization problem and derive a strong duality result, which enables us to provide model-based and model-free algorithms for SLQR optimal policy design based on the Karush-Kuhn-Tucker (KKT) conditions. An illustrative example demonstrates the proposed model-free algorithm's validity, showcasing the central nervous system's learning mechanism in human arm movement.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02613
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning-based primal-dual optimal control of discrete-time stochastic systems with multiplicative noise
Jiang, Xiushan
Zhang, Weihai
Optimization and Control
Reinforcement learning (RL) is an effective approach for solving optimal control problems without knowing the exact information of the system model. However, the classical Q-learning method, a model-free RL algorithm, has its limitations, such as lack of strict theoretical analysis and the need for artificial disturbances during implementation. This paper explores the partially model-free stochastic linear quadratic regulator (SLQR) problem for a system with multiplicative noise from the primal-dual perspective to address these challenges. This approach lays a strong theoretical foundation for understanding the intrinsic mechanisms of classical RL algorithms. We reformulate the SLQR into a non-convex primal-dual optimization problem and derive a strong duality result, which enables us to provide model-based and model-free algorithms for SLQR optimal policy design based on the Karush-Kuhn-Tucker (KKT) conditions. An illustrative example demonstrates the proposed model-free algorithm's validity, showcasing the central nervous system's learning mechanism in human arm movement.
title Learning-based primal-dual optimal control of discrete-time stochastic systems with multiplicative noise
topic Optimization and Control
url https://arxiv.org/abs/2506.02613