Repeat After Me: Transformers are Better than State Space Models at Copying

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jelassi, Samy, Brandfonbrener, David, Kakade, Sham M., Malach, Eran
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911901302128640
author Jelassi, Samy
Brandfonbrener, David
Kakade, Sham M.
Malach, Eran
author_facet Jelassi, Samy
Brandfonbrener, David
Kakade, Sham M.
Malach, Eran
contents Transformers are the dominant architecture for sequence modeling, but there is growing interest in models that use a fixed-size latent state that does not depend on the sequence length, which we refer to as "generalized state space models" (GSSMs). In this paper we show that while GSSMs are promising in terms of inference-time efficiency, they are limited compared to transformer models on tasks that require copying from the input context. We start with a theoretical analysis of the simple task of string copying and prove that a two layer transformer can copy strings of exponential length while GSSMs are fundamentally limited by their fixed-size latent state. Empirically, we find that transformers outperform GSSMs in terms of efficiency and generalization on synthetic tasks that require copying the context. Finally, we evaluate pretrained large language models and find that transformer models dramatically outperform state space models at copying and retrieving information from context. Taken together, these results suggest a fundamental gap between transformers and GSSMs on tasks of practical interest.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01032
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Repeat After Me: Transformers are Better than State Space Models at Copying
Jelassi, Samy
Brandfonbrener, David
Kakade, Sham M.
Malach, Eran
Machine Learning
Artificial Intelligence
Computation and Language
Transformers are the dominant architecture for sequence modeling, but there is growing interest in models that use a fixed-size latent state that does not depend on the sequence length, which we refer to as "generalized state space models" (GSSMs). In this paper we show that while GSSMs are promising in terms of inference-time efficiency, they are limited compared to transformer models on tasks that require copying from the input context. We start with a theoretical analysis of the simple task of string copying and prove that a two layer transformer can copy strings of exponential length while GSSMs are fundamentally limited by their fixed-size latent state. Empirically, we find that transformers outperform GSSMs in terms of efficiency and generalization on synthetic tasks that require copying the context. Finally, we evaluate pretrained large language models and find that transformer models dramatically outperform state space models at copying and retrieving information from context. Taken together, these results suggest a fundamental gap between transformers and GSSMs on tasks of practical interest.
title Repeat After Me: Transformers are Better than State Space Models at Copying
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2402.01032