Fixed-Point RNNs: Interpolating from Diagonal to Dense

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Movahedi, Sajad, Sarnthein, Felix, Cirone, Nicola Muca, Orvieto, Antonio
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909868117458944
author Movahedi, Sajad
Sarnthein, Felix
Cirone, Nicola Muca
Orvieto, Antonio
author_facet Movahedi, Sajad
Sarnthein, Felix
Cirone, Nicola Muca
Orvieto, Antonio
contents Linear recurrent neural networks (RNNs) and state-space models (SSMs) such as Mamba have become promising alternatives to softmax-attention as sequence mixing layers in Transformer architectures. Current models, however, do not exhibit the full state-tracking expressivity of RNNs because they rely on channel-wise (i.e. diagonal) sequence mixing. In this paper, we investigate parameterizations of a large class of dense linear RNNs as fixed-points of parallelizable diagonal linear RNNs. The resulting models can naturally trade expressivity for efficiency at a fixed number of parameters and achieve state-of-the-art results on the state-tracking benchmarks $A_5$ and $S_5$, while matching performance on copying and other tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10799
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fixed-Point RNNs: Interpolating from Diagonal to Dense
Movahedi, Sajad
Sarnthein, Felix
Cirone, Nicola Muca
Orvieto, Antonio
Machine Learning
Linear recurrent neural networks (RNNs) and state-space models (SSMs) such as Mamba have become promising alternatives to softmax-attention as sequence mixing layers in Transformer architectures. Current models, however, do not exhibit the full state-tracking expressivity of RNNs because they rely on channel-wise (i.e. diagonal) sequence mixing. In this paper, we investigate parameterizations of a large class of dense linear RNNs as fixed-points of parallelizable diagonal linear RNNs. The resulting models can naturally trade expressivity for efficiency at a fixed number of parameters and achieve state-of-the-art results on the state-tracking benchmarks $A_5$ and $S_5$, while matching performance on copying and other tasks.
title Fixed-Point RNNs: Interpolating from Diagonal to Dense
topic Machine Learning
url https://arxiv.org/abs/2503.10799