A Survey on Transformer Compression

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Yehui, Wang, Yunhe, Guo, Jianyuan, Tu, Zhijun, Han, Kai, Hu, Hailin, Tao, Dacheng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910401302626304
author Tang, Yehui
Wang, Yunhe
Guo, Jianyuan
Tu, Zhijun
Han, Kai
Hu, Hailin
Tao, Dacheng
author_facet Tang, Yehui
Wang, Yunhe
Guo, Jianyuan
Tu, Zhijun
Han, Kai
Hu, Hailin
Tao, Dacheng
contents Transformer plays a vital role in the realms of natural language processing (NLP) and computer vision (CV), specially for constructing large language models (LLM) and large vision models (LVM). Model compression methods reduce the memory and computational cost of Transformer, which is a necessary step to implement large language/vision models on practical devices. Given the unique architecture of Transformer, featuring alternative attention and feedforward neural network (FFN) modules, specific compression techniques are usually required. The efficiency of these compression methods is also paramount, as retraining large models on the entire training dataset is usually impractical. This survey provides a comprehensive review of recent compression methods, with a specific focus on their application to Transformer-based models. The compression methods are primarily categorized into pruning, quantization, knowledge distillation, and efficient architecture design (Mamba, RetNet, RWKV, etc.). In each category, we discuss compression methods for both language and vision tasks, highlighting common underlying principles. Finally, we delve into the relation between various compression methods, and discuss further directions in this domain.
format Preprint
id arxiv_https___arxiv_org_abs_2402_05964
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey on Transformer Compression
Tang, Yehui
Wang, Yunhe
Guo, Jianyuan
Tu, Zhijun
Han, Kai
Hu, Hailin
Tao, Dacheng
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Transformer plays a vital role in the realms of natural language processing (NLP) and computer vision (CV), specially for constructing large language models (LLM) and large vision models (LVM). Model compression methods reduce the memory and computational cost of Transformer, which is a necessary step to implement large language/vision models on practical devices. Given the unique architecture of Transformer, featuring alternative attention and feedforward neural network (FFN) modules, specific compression techniques are usually required. The efficiency of these compression methods is also paramount, as retraining large models on the entire training dataset is usually impractical. This survey provides a comprehensive review of recent compression methods, with a specific focus on their application to Transformer-based models. The compression methods are primarily categorized into pruning, quantization, knowledge distillation, and efficient architecture design (Mamba, RetNet, RWKV, etc.). In each category, we discuss compression methods for both language and vision tasks, highlighting common underlying principles. Finally, we delve into the relation between various compression methods, and discuss further directions in this domain.
title A Survey on Transformer Compression
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.05964