Long-Context Linear System Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yüksel, Oğuz Kaan, Even, Mathieu, Flammarion, Nicolas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908430239793152
author Yüksel, Oğuz Kaan
Even, Mathieu
Flammarion, Nicolas
author_facet Yüksel, Oğuz Kaan
Even, Mathieu
Flammarion, Nicolas
contents This paper addresses the problem of long-context linear system identification, where the state $x_t$ of a dynamical system at time $t$ depends linearly on previous states $x_s$ over a fixed context window of length $p$. We establish a sample complexity bound that matches the i.i.d. parametric rate up to logarithmic factors for a broad class of systems, extending previous works that considered only first-order dependencies. Our findings reveal a learning-without-mixing phenomenon, indicating that learning long-context linear autoregressive models is not hindered by slow mixing properties potentially associated with extended context windows. Additionally, we extend these results to (i) shared low-rank representations, where rank-regularized estimators improve the dependence of the rates on the dimensionality, and (ii) misspecified context lengths in strictly stable systems, where shorter contexts offer statistical advantages.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05690
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Long-Context Linear System Identification
Yüksel, Oğuz Kaan
Even, Mathieu
Flammarion, Nicolas
Machine Learning
Systems and Control
Statistics Theory
This paper addresses the problem of long-context linear system identification, where the state $x_t$ of a dynamical system at time $t$ depends linearly on previous states $x_s$ over a fixed context window of length $p$. We establish a sample complexity bound that matches the i.i.d. parametric rate up to logarithmic factors for a broad class of systems, extending previous works that considered only first-order dependencies. Our findings reveal a learning-without-mixing phenomenon, indicating that learning long-context linear autoregressive models is not hindered by slow mixing properties potentially associated with extended context windows. Additionally, we extend these results to (i) shared low-rank representations, where rank-regularized estimators improve the dependence of the rates on the dimensionality, and (ii) misspecified context lengths in strictly stable systems, where shorter contexts offer statistical advantages.
title Long-Context Linear System Identification
topic Machine Learning
Systems and Control
Statistics Theory
url https://arxiv.org/abs/2410.05690