State-space models can learn in-context by gradient descent

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sushma, Neeraj Mohan, Tian, Yudou, Mestha, Harshvardhan, Colombo, Nicolo, Kappel, David, Subramoney, Anand
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909500342009856
author Sushma, Neeraj Mohan
Tian, Yudou
Mestha, Harshvardhan
Colombo, Nicolo
Kappel, David
Subramoney, Anand
author_facet Sushma, Neeraj Mohan
Tian, Yudou
Mestha, Harshvardhan
Colombo, Nicolo
Kappel, David
Subramoney, Anand
contents Deep state-space models (Deep SSMs) are becoming popular as effective approaches to model sequence data. They have also been shown to be capable of in-context learning, much like transformers. However, a complete picture of how SSMs might be able to do in-context learning has been missing. In this study, we provide a direct and explicit construction to show that state-space models can perform gradient-based learning and use it for in-context learning in much the same way as transformers. Specifically, we prove that a single structured state-space model layer, augmented with multiplicative input and output gating, can reproduce the outputs of an implicit linear model with least squares loss after one step of gradient descent. We then show a straightforward extension to multi-step linear and non-linear regression tasks. We validate our construction by training randomly initialized augmented SSMs on linear and non-linear regression tasks. The empirically obtained parameters through optimization match the ones predicted analytically by the theoretical construction. Overall, we elucidate the role of input- and output-gating in recurrent architectures as the key inductive biases for enabling the expressive power typical of foundation models. We also provide novel insights into the relationship between state-space models and linear self-attention, and their ability to learn in-context.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11687
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle State-space models can learn in-context by gradient descent
Sushma, Neeraj Mohan
Tian, Yudou
Mestha, Harshvardhan
Colombo, Nicolo
Kappel, David
Subramoney, Anand
Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
Deep state-space models (Deep SSMs) are becoming popular as effective approaches to model sequence data. They have also been shown to be capable of in-context learning, much like transformers. However, a complete picture of how SSMs might be able to do in-context learning has been missing. In this study, we provide a direct and explicit construction to show that state-space models can perform gradient-based learning and use it for in-context learning in much the same way as transformers. Specifically, we prove that a single structured state-space model layer, augmented with multiplicative input and output gating, can reproduce the outputs of an implicit linear model with least squares loss after one step of gradient descent. We then show a straightforward extension to multi-step linear and non-linear regression tasks. We validate our construction by training randomly initialized augmented SSMs on linear and non-linear regression tasks. The empirically obtained parameters through optimization match the ones predicted analytically by the theoretical construction. Overall, we elucidate the role of input- and output-gating in recurrent architectures as the key inductive biases for enabling the expressive power typical of foundation models. We also provide novel insights into the relationship between state-space models and linear self-attention, and their ability to learn in-context.
title State-space models can learn in-context by gradient descent
topic Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2410.11687