On the Robustness of Derivative-free Methods for Linear Quadratic Regulator

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Weijian, Kounatidis, Panagiotis, Jiang, Zhong-Ping, Malikopoulos, Andreas A.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909649230364672
author Li, Weijian
Kounatidis, Panagiotis
Jiang, Zhong-Ping
Malikopoulos, Andreas A.
author_facet Li, Weijian
Kounatidis, Panagiotis
Jiang, Zhong-Ping
Malikopoulos, Andreas A.
contents Policy optimization has drawn increasing attention in reinforcement learning, particularly in the context of derivative-free methods for linear quadratic regulator (LQR) problems with unknown dynamics. This paper focuses on characterizing the robustness of derivative-free methods for solving an infinite-horizon LQR problem. To be specific, we estimate policy gradients by cost values, and study the effect of perturbations on the estimations, where the perturbations may arise from function approximations, measurement noises, etc. We show that under sufficiently small perturbations, the derivative-free methods converge to any pre-specified neighborhood of the optimal policy. Furthermore, we establish explicit bounds on the perturbations, and provide the sample complexity for the perturbed derivative-free methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12596
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the Robustness of Derivative-free Methods for Linear Quadratic Regulator
Li, Weijian
Kounatidis, Panagiotis
Jiang, Zhong-Ping
Malikopoulos, Andreas A.
Optimization and Control
Policy optimization has drawn increasing attention in reinforcement learning, particularly in the context of derivative-free methods for linear quadratic regulator (LQR) problems with unknown dynamics. This paper focuses on characterizing the robustness of derivative-free methods for solving an infinite-horizon LQR problem. To be specific, we estimate policy gradients by cost values, and study the effect of perturbations on the estimations, where the perturbations may arise from function approximations, measurement noises, etc. We show that under sufficiently small perturbations, the derivative-free methods converge to any pre-specified neighborhood of the optimal policy. Furthermore, we establish explicit bounds on the perturbations, and provide the sample complexity for the perturbed derivative-free methods.
title On the Robustness of Derivative-free Methods for Linear Quadratic Regulator
topic Optimization and Control
url https://arxiv.org/abs/2506.12596