Regret Lower Bounds for Learning Linear Quadratic Gaussian Systems

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ziemann, Ingvar, Sandberg, Henrik
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917690793263104
author Ziemann, Ingvar
Sandberg, Henrik
author_facet Ziemann, Ingvar
Sandberg, Henrik
contents TWe establish regret lower bounds for adaptively controlling an unknown linear Gaussian system with quadratic costs. We combine ideas from experiment design, estimation theory and a perturbation bound of certain information matrices to derive regret lower bounds exhibiting scaling on the order of magnitude $\sqrt{T}$ in the time horizon $T$. Our bounds accurately capture the role of control-theoretic parameters and we are able to show that systems that are hard to control are also hard to learn to control; when instantiated to state feedback systems we recover the dimensional dependency of earlier work but with improved scaling with system-theoretic constants such as system costs and Gramians. Furthermore, we extend our results to a class of partially observed systems and demonstrate that systems with poor observability structure also are hard to learn to control.
format Preprint
id arxiv_https___arxiv_org_abs_2201_01680
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Regret Lower Bounds for Learning Linear Quadratic Gaussian Systems
Ziemann, Ingvar
Sandberg, Henrik
Machine Learning
Optimization and Control
TWe establish regret lower bounds for adaptively controlling an unknown linear Gaussian system with quadratic costs. We combine ideas from experiment design, estimation theory and a perturbation bound of certain information matrices to derive regret lower bounds exhibiting scaling on the order of magnitude $\sqrt{T}$ in the time horizon $T$. Our bounds accurately capture the role of control-theoretic parameters and we are able to show that systems that are hard to control are also hard to learn to control; when instantiated to state feedback systems we recover the dimensional dependency of earlier work but with improved scaling with system-theoretic constants such as system costs and Gramians. Furthermore, we extend our results to a class of partially observed systems and demonstrate that systems with poor observability structure also are hard to learn to control.
title Regret Lower Bounds for Learning Linear Quadratic Gaussian Systems
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2201.01680