Policy Gradient Bounds in Multitask LQR

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Stamouli, Charis, Toso, Leonardo F., Tsiamis, Anastasios, Pappas, George J., Anderson, James
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916964965810176
author Stamouli, Charis
Toso, Leonardo F.
Tsiamis, Anastasios
Pappas, George J.
Anderson, James
author_facet Stamouli, Charis
Toso, Leonardo F.
Tsiamis, Anastasios
Pappas, George J.
Anderson, James
contents We analyze the performance of policy gradient in multitask linear quadratic regulation (LQR), where the system and cost parameters differ across tasks. The main goal of multitask LQR is to find a controller with satisfactory performance on every task. Prior analyses on relevant contexts fail to capture closed-loop task similarities, resulting in conservative performance guarantees. To account for such similarities, we propose bisimulation-based measures of task heterogeneity. Our measures employ new bisimulation functions to bound the cost gradient distance between a pair of tasks in closed loop with a common stabilizing controller. Employing these measures, we derive suboptimality bounds for both the multitask optimal controller and the asymptotic policy gradient controller with respect to each of the tasks. We further provide conditions under which the policy gradient iterates remain stabilizing for every system. For multiple random sets of certain tasks, we observe that our bisimulation-based measures improve upon baseline measures of task heterogeneity dramatically.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19266
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Policy Gradient Bounds in Multitask LQR
Stamouli, Charis
Toso, Leonardo F.
Tsiamis, Anastasios
Pappas, George J.
Anderson, James
Systems and Control
Optimization and Control
We analyze the performance of policy gradient in multitask linear quadratic regulation (LQR), where the system and cost parameters differ across tasks. The main goal of multitask LQR is to find a controller with satisfactory performance on every task. Prior analyses on relevant contexts fail to capture closed-loop task similarities, resulting in conservative performance guarantees. To account for such similarities, we propose bisimulation-based measures of task heterogeneity. Our measures employ new bisimulation functions to bound the cost gradient distance between a pair of tasks in closed loop with a common stabilizing controller. Employing these measures, we derive suboptimality bounds for both the multitask optimal controller and the asymptotic policy gradient controller with respect to each of the tasks. We further provide conditions under which the policy gradient iterates remain stabilizing for every system. For multiple random sets of certain tasks, we observe that our bisimulation-based measures improve upon baseline measures of task heterogeneity dramatically.
title Policy Gradient Bounds in Multitask LQR
topic Systems and Control
Optimization and Control
url https://arxiv.org/abs/2509.19266