Enhanced Transformer architecture for in-context learning of dynamical systems

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rufolo, Matteo, Piga, Dario, Maroni, Gabriele, Forgione, Marco
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916423288225792
author Rufolo, Matteo
Piga, Dario
Maroni, Gabriele
Forgione, Marco
author_facet Rufolo, Matteo
Piga, Dario
Maroni, Gabriele
Forgione, Marco
contents Recently introduced by some of the authors, the in-context identification paradigm aims at estimating, offline and based on synthetic data, a meta-model that describes the behavior of a whole class of systems. Once trained, this meta-model is fed with an observed input/output sequence (context) generated by a real system to predict its behavior in a zero-shot learning fashion. In this paper, we enhance the original meta-modeling framework through three key innovations: by formulating the learning task within a probabilistic framework; by managing non-contiguous context and query windows; and by adopting recurrent patching to effectively handle long context sequences. The efficacy of these modifications is demonstrated through a numerical example focusing on the Wiener-Hammerstein system class, highlighting the model's enhanced performance and scalability.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03291
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhanced Transformer architecture for in-context learning of dynamical systems
Rufolo, Matteo
Piga, Dario
Maroni, Gabriele
Forgione, Marco
Machine Learning
Artificial Intelligence
Systems and Control
Recently introduced by some of the authors, the in-context identification paradigm aims at estimating, offline and based on synthetic data, a meta-model that describes the behavior of a whole class of systems. Once trained, this meta-model is fed with an observed input/output sequence (context) generated by a real system to predict its behavior in a zero-shot learning fashion. In this paper, we enhance the original meta-modeling framework through three key innovations: by formulating the learning task within a probabilistic framework; by managing non-contiguous context and query windows; and by adopting recurrent patching to effectively handle long context sequences. The efficacy of these modifications is demonstrated through a numerical example focusing on the Wiener-Hammerstein system class, highlighting the model's enhanced performance and scalability.
title Enhanced Transformer architecture for in-context learning of dynamical systems
topic Machine Learning
Artificial Intelligence
Systems and Control
url https://arxiv.org/abs/2410.03291