Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: De, Soham, Smith, Samuel L., Fernando, Anushan, Botev, Aleksandar, Cristian-Muraru, George, Gu, Albert, Haroun, Ruba, Berrada, Leonard, Chen, Yutian, Srinivasan, Srivatsan, Desjardins, Guillaume, Doucet, Arnaud, Budden, David, Teh, Yee Whye, Pascanu, Razvan, De Freitas, Nando, Gulcehre, Caglar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914697183232000
author De, Soham
Smith, Samuel L.
Fernando, Anushan
Botev, Aleksandar
Cristian-Muraru, George
Gu, Albert
Haroun, Ruba
Berrada, Leonard
Chen, Yutian
Srinivasan, Srivatsan
Desjardins, Guillaume
Doucet, Arnaud
Budden, David
Teh, Yee Whye
Pascanu, Razvan
De Freitas, Nando
Gulcehre, Caglar
author_facet De, Soham
Smith, Samuel L.
Fernando, Anushan
Botev, Aleksandar
Cristian-Muraru, George
Gu, Albert
Haroun, Ruba
Berrada, Leonard
Chen, Yutian
Srinivasan, Srivatsan
Desjardins, Guillaume
Doucet, Arnaud
Budden, David
Teh, Yee Whye
Pascanu, Razvan
De Freitas, Nando
Gulcehre, Caglar
contents Recurrent neural networks (RNNs) have fast inference and scale efficiently on long sequences, but they are difficult to train and hard to scale. We propose Hawk, an RNN with gated linear recurrences, and Griffin, a hybrid model that mixes gated linear recurrences with local attention. Hawk exceeds the reported performance of Mamba on downstream tasks, while Griffin matches the performance of Llama-2 despite being trained on over 6 times fewer tokens. We also show that Griffin can extrapolate on sequences significantly longer than those seen during training. Our models match the hardware efficiency of Transformers during training, and during inference they have lower latency and significantly higher throughput. We scale Griffin up to 14B parameters, and explain how to shard our models for efficient distributed training.
format Preprint
id arxiv_https___arxiv_org_abs_2402_19427
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
De, Soham
Smith, Samuel L.
Fernando, Anushan
Botev, Aleksandar
Cristian-Muraru, George
Gu, Albert
Haroun, Ruba
Berrada, Leonard
Chen, Yutian
Srinivasan, Srivatsan
Desjardins, Guillaume
Doucet, Arnaud
Budden, David
Teh, Yee Whye
Pascanu, Razvan
De Freitas, Nando
Gulcehre, Caglar
Machine Learning
Computation and Language
Recurrent neural networks (RNNs) have fast inference and scale efficiently on long sequences, but they are difficult to train and hard to scale. We propose Hawk, an RNN with gated linear recurrences, and Griffin, a hybrid model that mixes gated linear recurrences with local attention. Hawk exceeds the reported performance of Mamba on downstream tasks, while Griffin matches the performance of Llama-2 despite being trained on over 6 times fewer tokens. We also show that Griffin can extrapolate on sequences significantly longer than those seen during training. Our models match the hardware efficiency of Transformers during training, and during inference they have lower latency and significantly higher throughput. We scale Griffin up to 14B parameters, and explain how to shard our models for efficient distributed training.
title Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2402.19427