A Stochastic Gradient Descent Approach to Design Policy Gradient Methods for LQR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Bowen, Weissmann, Simon, Staudigl, Mathias, Iannelli, Andrea
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912917559967744
author Song, Bowen
Weissmann, Simon
Staudigl, Mathias
Iannelli, Andrea
author_facet Song, Bowen
Weissmann, Simon
Staudigl, Mathias
Iannelli, Andrea
contents In this work, we propose a stochastic gradient descent (SGD) framework to design data-driven policy gradient descent algorithms for the linear quadratic regulator problem. Two alternative schemes are considered to estimate the policy gradient from stochastic trajectory data: (i) an indirect online identification based approach, in which the system matrices are first estimated and subsequently used to construct the gradient, and (ii) a direct zeroth-order approach, which approximates the gradient using empirical cost evaluations. In both cases, the resulting gradient estimates are random due to stochasticity in the data, allowing us to use SGD theory to analyze the convergence of the associated policy gradient methods. A key technical step consists of modeling the gradient estimates as suitable stochastic gradient oracles, which, because of the way they are computed, are inherently based. We derive sufficient conditions under which SGD with a biased gradient oracle converges asymptotically to the optimal policy, and leverage these conditions to design the parameters of the gradient estimation schemes. Moreover, we compare the advantages and limitations of the two data-driven gradient estimators. Numerical experiments validate the effectiveness of the proposed methods.
format Preprint
id arxiv_https___arxiv_org_abs_2602_18933
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Stochastic Gradient Descent Approach to Design Policy Gradient Methods for LQR
Song, Bowen
Weissmann, Simon
Staudigl, Mathias
Iannelli, Andrea
Systems and Control
In this work, we propose a stochastic gradient descent (SGD) framework to design data-driven policy gradient descent algorithms for the linear quadratic regulator problem. Two alternative schemes are considered to estimate the policy gradient from stochastic trajectory data: (i) an indirect online identification based approach, in which the system matrices are first estimated and subsequently used to construct the gradient, and (ii) a direct zeroth-order approach, which approximates the gradient using empirical cost evaluations. In both cases, the resulting gradient estimates are random due to stochasticity in the data, allowing us to use SGD theory to analyze the convergence of the associated policy gradient methods. A key technical step consists of modeling the gradient estimates as suitable stochastic gradient oracles, which, because of the way they are computed, are inherently based. We derive sufficient conditions under which SGD with a biased gradient oracle converges asymptotically to the optimal policy, and leverage these conditions to design the parameters of the gradient estimation schemes. Moreover, we compare the advantages and limitations of the two data-driven gradient estimators. Numerical experiments validate the effectiveness of the proposed methods.
title A Stochastic Gradient Descent Approach to Design Policy Gradient Methods for LQR
topic Systems and Control
url https://arxiv.org/abs/2602.18933