Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dong, Pingcheng, Tan, Yonghao, Zhang, Dong, Ni, Tianwei, Liu, Xuejiao, Liu, Yu, Luo, Peng, Liang, Luhong, Liu, Shih-Yang, Huang, Xijie, Zhu, Huaiyu, Pan, Yun, An, Fengwei, Cheng, Kwang-Ting
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917625720733696
author Dong, Pingcheng
Tan, Yonghao
Zhang, Dong
Ni, Tianwei
Liu, Xuejiao
Liu, Yu
Luo, Peng
Liang, Luhong
Liu, Shih-Yang
Huang, Xijie
Zhu, Huaiyu
Pan, Yun
An, Fengwei
Cheng, Kwang-Ting
author_facet Dong, Pingcheng
Tan, Yonghao
Zhang, Dong
Ni, Tianwei
Liu, Xuejiao
Liu, Yu
Luo, Peng
Liang, Luhong
Liu, Shih-Yang
Huang, Xijie
Zhu, Huaiyu
Pan, Yun
An, Fengwei
Cheng, Kwang-Ting
contents Non-linear functions are prevalent in Transformers and their lightweight variants, incurring substantial and frequently underestimated hardware costs. Previous state-of-the-art works optimize these operations by piece-wise linear approximation and store the parameters in look-up tables (LUT), but most of them require unfriendly high-precision arithmetics such as FP/INT 32 and lack consideration of integer-only INT quantization. This paper proposed a genetic LUT-Approximation algorithm namely GQA-LUT that can automatically determine the parameters with quantization awareness. The results demonstrate that GQA-LUT achieves negligible degradation on the challenging semantic segmentation task for both vanilla and linear Transformer models. Besides, proposed GQA-LUT enables the employment of INT8-based LUT-Approximation that achieves an area savings of 81.3~81.7% and a power reduction of 79.3~80.2% compared to the high-precision FP/INT 32 alternatives. Code is available at https:// github.com/PingchengDong/GQA-LUT.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19591
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
Dong, Pingcheng
Tan, Yonghao
Zhang, Dong
Ni, Tianwei
Liu, Xuejiao
Liu, Yu
Luo, Peng
Liang, Luhong
Liu, Shih-Yang
Huang, Xijie
Zhu, Huaiyu
Pan, Yun
An, Fengwei
Cheng, Kwang-Ting
Machine Learning
Hardware Architecture
Neural and Evolutionary Computing
Non-linear functions are prevalent in Transformers and their lightweight variants, incurring substantial and frequently underestimated hardware costs. Previous state-of-the-art works optimize these operations by piece-wise linear approximation and store the parameters in look-up tables (LUT), but most of them require unfriendly high-precision arithmetics such as FP/INT 32 and lack consideration of integer-only INT quantization. This paper proposed a genetic LUT-Approximation algorithm namely GQA-LUT that can automatically determine the parameters with quantization awareness. The results demonstrate that GQA-LUT achieves negligible degradation on the challenging semantic segmentation task for both vanilla and linear Transformer models. Besides, proposed GQA-LUT enables the employment of INT8-based LUT-Approximation that achieves an area savings of 81.3~81.7% and a power reduction of 79.3~80.2% compared to the high-precision FP/INT 32 alternatives. Code is available at https:// github.com/PingchengDong/GQA-LUT.
title Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers
topic Machine Learning
Hardware Architecture
Neural and Evolutionary Computing
url https://arxiv.org/abs/2403.19591