BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ji, Yuhao, Fang, Chao, Wang, Zhongfeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914868470218752
author Ji, Yuhao
Fang, Chao
Wang, Zhongfeng
author_facet Ji, Yuhao
Fang, Chao
Wang, Zhongfeng
contents Existing binary Transformers are promising in edge deployment due to their compact model size, low computational complexity, and considerable inference accuracy. However, deploying binary Transformers faces challenges on prior processors due to inefficient execution of quantized matrix multiplication (QMM) and the energy consumption overhead caused by multi-precision activations. To tackle the challenges above, we first develop a computation flow abstraction method for binary Transformers to improve QMM execution efficiency by optimizing the computation order. Furthermore, a binarized energy-efficient Transformer accelerator, namely BETA, is proposed to boost the efficient deployment at the edge. Notably, BETA features a configurable QMM engine, accommodating diverse activation precisions of binary Transformers and offering high-parallelism and high-speed for QMMs with impressive energy efficiency. Experimental results evaluated on ZCU102 FPGA show BETA achieves an average energy efficiency of 174 GOPS/W, which is 1.76~21.92x higher than prior FPGA-based accelerators, showing BETA's good potential for edge Transformer acceleration.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
Ji, Yuhao
Fang, Chao
Wang, Zhongfeng
Hardware Architecture
Artificial Intelligence
Existing binary Transformers are promising in edge deployment due to their compact model size, low computational complexity, and considerable inference accuracy. However, deploying binary Transformers faces challenges on prior processors due to inefficient execution of quantized matrix multiplication (QMM) and the energy consumption overhead caused by multi-precision activations. To tackle the challenges above, we first develop a computation flow abstraction method for binary Transformers to improve QMM execution efficiency by optimizing the computation order. Furthermore, a binarized energy-efficient Transformer accelerator, namely BETA, is proposed to boost the efficient deployment at the edge. Notably, BETA features a configurable QMM engine, accommodating diverse activation precisions of binary Transformers and offering high-parallelism and high-speed for QMMs with impressive energy efficiency. Experimental results evaluated on ZCU102 FPGA show BETA achieves an average energy efficiency of 174 GOPS/W, which is 1.76~21.92x higher than prior FPGA-based accelerators, showing BETA's good potential for edge Transformer acceleration.
title BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
topic Hardware Architecture
Artificial Intelligence
url https://arxiv.org/abs/2401.11851