Revenge of the Fallen? Recurrent Models Match Transformers at Predicting Human Language Comprehension Metrics

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Michaelov, James A., Arnett, Catherine, Bergen, Benjamin K.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916369211064320
author Michaelov, James A.
Arnett, Catherine
Bergen, Benjamin K.
author_facet Michaelov, James A.
Arnett, Catherine
Bergen, Benjamin K.
contents Transformers have generally supplanted recurrent neural networks as the dominant architecture for both natural language processing tasks and for modelling the effect of predictability on online human language comprehension. However, two recently developed recurrent model architectures, RWKV and Mamba, appear to perform natural language tasks comparably to or better than transformers of equivalent scale. In this paper, we show that contemporary recurrent models are now also able to match - and in some cases, exceed - the performance of comparably sized transformers at modeling online human language comprehension. This suggests that transformer language models are not uniquely suited to this task, and opens up new directions for debates about the extent to which architectural features of language models make them better or worse models of human language comprehension.
format Preprint
id arxiv_https___arxiv_org_abs_2404_19178
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revenge of the Fallen? Recurrent Models Match Transformers at Predicting Human Language Comprehension Metrics
Michaelov, James A.
Arnett, Catherine
Bergen, Benjamin K.
Computation and Language
Transformers have generally supplanted recurrent neural networks as the dominant architecture for both natural language processing tasks and for modelling the effect of predictability on online human language comprehension. However, two recently developed recurrent model architectures, RWKV and Mamba, appear to perform natural language tasks comparably to or better than transformers of equivalent scale. In this paper, we show that contemporary recurrent models are now also able to match - and in some cases, exceed - the performance of comparably sized transformers at modeling online human language comprehension. This suggests that transformer language models are not uniquely suited to this task, and opens up new directions for debates about the extent to which architectural features of language models make them better or worse models of human language comprehension.
title Revenge of the Fallen? Recurrent Models Match Transformers at Predicting Human Language Comprehension Metrics
topic Computation and Language
url https://arxiv.org/abs/2404.19178