IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929335614570496 |
|---|---|
| author | Mao, Yuzhen Ester, Martin Li, Ke |
| author_facet | Mao, Yuzhen Ester, Martin Li, Ke |
| contents | One limitation of existing Transformer-based models is that they cannot handle very long sequences as input since their self-attention operations exhibit quadratic time and space complexity. This problem becomes especially acute when Transformers are deployed on hardware platforms equipped only with CPUs. To address this issue, we propose a novel method for accelerating self-attention at inference time that works with pretrained Transformer models out-of-the-box without requiring retraining. We experiment using our method to accelerate various long-sequence Transformers, including a leading LLaMA 2-based LLM, on various benchmarks and demonstrate a greater speedup of 2.73x - 7.63x while retaining 98.6% - 99.6% of the accuracy of the original pretrained models. The code is available on our project website at https://yuzhenmao.github.io/IceFormer/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_02842 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs Mao, Yuzhen Ester, Martin Li, Ke Machine Learning One limitation of existing Transformer-based models is that they cannot handle very long sequences as input since their self-attention operations exhibit quadratic time and space complexity. This problem becomes especially acute when Transformers are deployed on hardware platforms equipped only with CPUs. To address this issue, we propose a novel method for accelerating self-attention at inference time that works with pretrained Transformer models out-of-the-box without requiring retraining. We experiment using our method to accelerate various long-sequence Transformers, including a leading LLaMA 2-based LLM, on various benchmarks and demonstrate a greater speedup of 2.73x - 7.63x while retaining 98.6% - 99.6% of the accuracy of the original pretrained models. The code is available on our project website at https://yuzhenmao.github.io/IceFormer/. |
| title | IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2405.02842 |