Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914697183232000 |
|---|---|
| author | De, Soham Smith, Samuel L. Fernando, Anushan Botev, Aleksandar Cristian-Muraru, George Gu, Albert Haroun, Ruba Berrada, Leonard Chen, Yutian Srinivasan, Srivatsan Desjardins, Guillaume Doucet, Arnaud Budden, David Teh, Yee Whye Pascanu, Razvan De Freitas, Nando Gulcehre, Caglar |
| author_facet | De, Soham Smith, Samuel L. Fernando, Anushan Botev, Aleksandar Cristian-Muraru, George Gu, Albert Haroun, Ruba Berrada, Leonard Chen, Yutian Srinivasan, Srivatsan Desjardins, Guillaume Doucet, Arnaud Budden, David Teh, Yee Whye Pascanu, Razvan De Freitas, Nando Gulcehre, Caglar |
| contents | Recurrent neural networks (RNNs) have fast inference and scale efficiently on long sequences, but they are difficult to train and hard to scale. We propose Hawk, an RNN with gated linear recurrences, and Griffin, a hybrid model that mixes gated linear recurrences with local attention. Hawk exceeds the reported performance of Mamba on downstream tasks, while Griffin matches the performance of Llama-2 despite being trained on over 6 times fewer tokens. We also show that Griffin can extrapolate on sequences significantly longer than those seen during training. Our models match the hardware efficiency of Transformers during training, and during inference they have lower latency and significantly higher throughput. We scale Griffin up to 14B parameters, and explain how to shard our models for efficient distributed training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_19427 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models De, Soham Smith, Samuel L. Fernando, Anushan Botev, Aleksandar Cristian-Muraru, George Gu, Albert Haroun, Ruba Berrada, Leonard Chen, Yutian Srinivasan, Srivatsan Desjardins, Guillaume Doucet, Arnaud Budden, David Teh, Yee Whye Pascanu, Razvan De Freitas, Nando Gulcehre, Caglar Machine Learning Computation and Language Recurrent neural networks (RNNs) have fast inference and scale efficiently on long sequences, but they are difficult to train and hard to scale. We propose Hawk, an RNN with gated linear recurrences, and Griffin, a hybrid model that mixes gated linear recurrences with local attention. Hawk exceeds the reported performance of Mamba on downstream tasks, while Griffin matches the performance of Llama-2 despite being trained on over 6 times fewer tokens. We also show that Griffin can extrapolate on sequences significantly longer than those seen during training. Our models match the hardware efficiency of Transformers during training, and during inference they have lower latency and significantly higher throughput. We scale Griffin up to 14B parameters, and explain how to shard our models for efficient distributed training. |
| title | Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2402.19427 |