IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mao, Yuzhen, Ester, Martin, Li, Ke
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929335614570496
author Mao, Yuzhen
Ester, Martin
Li, Ke
author_facet Mao, Yuzhen
Ester, Martin
Li, Ke
contents One limitation of existing Transformer-based models is that they cannot handle very long sequences as input since their self-attention operations exhibit quadratic time and space complexity. This problem becomes especially acute when Transformers are deployed on hardware platforms equipped only with CPUs. To address this issue, we propose a novel method for accelerating self-attention at inference time that works with pretrained Transformer models out-of-the-box without requiring retraining. We experiment using our method to accelerate various long-sequence Transformers, including a leading LLaMA 2-based LLM, on various benchmarks and demonstrate a greater speedup of 2.73x - 7.63x while retaining 98.6% - 99.6% of the accuracy of the original pretrained models. The code is available on our project website at https://yuzhenmao.github.io/IceFormer/.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02842
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs
Mao, Yuzhen
Ester, Martin
Li, Ke
Machine Learning
One limitation of existing Transformer-based models is that they cannot handle very long sequences as input since their self-attention operations exhibit quadratic time and space complexity. This problem becomes especially acute when Transformers are deployed on hardware platforms equipped only with CPUs. To address this issue, we propose a novel method for accelerating self-attention at inference time that works with pretrained Transformer models out-of-the-box without requiring retraining. We experiment using our method to accelerate various long-sequence Transformers, including a leading LLaMA 2-based LLM, on various benchmarks and demonstrate a greater speedup of 2.73x - 7.63x while retaining 98.6% - 99.6% of the accuracy of the original pretrained models. The code is available on our project website at https://yuzhenmao.github.io/IceFormer/.
title IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs
topic Machine Learning
url https://arxiv.org/abs/2405.02842