Fixed-Point RNNs: Interpolating from Diagonal to Dense
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909868117458944 |
|---|---|
| author | Movahedi, Sajad Sarnthein, Felix Cirone, Nicola Muca Orvieto, Antonio |
| author_facet | Movahedi, Sajad Sarnthein, Felix Cirone, Nicola Muca Orvieto, Antonio |
| contents | Linear recurrent neural networks (RNNs) and state-space models (SSMs) such as Mamba have become promising alternatives to softmax-attention as sequence mixing layers in Transformer architectures. Current models, however, do not exhibit the full state-tracking expressivity of RNNs because they rely on channel-wise (i.e. diagonal) sequence mixing. In this paper, we investigate parameterizations of a large class of dense linear RNNs as fixed-points of parallelizable diagonal linear RNNs. The resulting models can naturally trade expressivity for efficiency at a fixed number of parameters and achieve state-of-the-art results on the state-tracking benchmarks $A_5$ and $S_5$, while matching performance on copying and other tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_10799 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Fixed-Point RNNs: Interpolating from Diagonal to Dense Movahedi, Sajad Sarnthein, Felix Cirone, Nicola Muca Orvieto, Antonio Machine Learning Linear recurrent neural networks (RNNs) and state-space models (SSMs) such as Mamba have become promising alternatives to softmax-attention as sequence mixing layers in Transformer architectures. Current models, however, do not exhibit the full state-tracking expressivity of RNNs because they rely on channel-wise (i.e. diagonal) sequence mixing. In this paper, we investigate parameterizations of a large class of dense linear RNNs as fixed-points of parallelizable diagonal linear RNNs. The resulting models can naturally trade expressivity for efficiency at a fixed number of parameters and achieve state-of-the-art results on the state-tracking benchmarks $A_5$ and $S_5$, while matching performance on copying and other tasks. |
| title | Fixed-Point RNNs: Interpolating from Diagonal to Dense |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2503.10799 |