SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914960952524800 |
|---|---|
| author | Csordás, Róbert Piękos, Piotr Irie, Kazuki Schmidhuber, Jürgen |
| author_facet | Csordás, Róbert Piękos, Piotr Irie, Kazuki Schmidhuber, Jürgen |
| contents | Despite many recent works on Mixture of Experts (MoEs) for resource-efficient Transformer language models, existing methods mostly focus on MoEs for feedforward layers. Previous attempts at extending MoE to the self-attention layer fail to match the performance of the parameter-matched baseline. Our novel SwitchHead is an effective MoE method for the attention layer that successfully reduces both the compute and memory requirements, achieving wall-clock speedup, while matching the language modeling performance of the baseline Transformer. Our novel MoE mechanism allows SwitchHead to compute up to 8 times fewer attention matrices than the standard Transformer. SwitchHead can also be combined with MoE feedforward layers, resulting in fully-MoE "SwitchAll" Transformers. For our 262M parameter model trained on C4, SwitchHead matches the perplexity of standard models with only 44% compute and 27% memory usage. Zero-shot experiments on downstream tasks confirm the performance of SwitchHead, e.g., achieving more than 3.5% absolute improvements on BliMP compared to the baseline with an equal compute resource. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_07987 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention Csordás, Róbert Piękos, Piotr Irie, Kazuki Schmidhuber, Jürgen Machine Learning Computation and Language Neural and Evolutionary Computing Despite many recent works on Mixture of Experts (MoEs) for resource-efficient Transformer language models, existing methods mostly focus on MoEs for feedforward layers. Previous attempts at extending MoE to the self-attention layer fail to match the performance of the parameter-matched baseline. Our novel SwitchHead is an effective MoE method for the attention layer that successfully reduces both the compute and memory requirements, achieving wall-clock speedup, while matching the language modeling performance of the baseline Transformer. Our novel MoE mechanism allows SwitchHead to compute up to 8 times fewer attention matrices than the standard Transformer. SwitchHead can also be combined with MoE feedforward layers, resulting in fully-MoE "SwitchAll" Transformers. For our 262M parameter model trained on C4, SwitchHead matches the perplexity of standard models with only 44% compute and 27% memory usage. Zero-shot experiments on downstream tasks confirm the performance of SwitchHead, e.g., achieving more than 3.5% absolute improvements on BliMP compared to the baseline with an equal compute resource. |
| title | SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention |
| topic | Machine Learning Computation and Language Neural and Evolutionary Computing |
| url | https://arxiv.org/abs/2312.07987 |