COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qiao, Ye, Chen, Zhiheng, Wang, Yian, Zhang, Yifan, Deng, Yunzhe, Huang, Sitao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912346258014208
author Qiao, Ye
Chen, Zhiheng
Wang, Yian
Zhang, Yifan
Deng, Yunzhe
Huang, Sitao
author_facet Qiao, Ye
Chen, Zhiheng
Wang, Yian
Zhang, Yifan
Deng, Yunzhe
Huang, Sitao
contents Transformer-based models have demonstrated superior performance in various fields, including natural language processing and computer vision. However, their enormous model size and high demands in computation, memory, and communication limit their deployment to edge platforms for local, secure inference. Binary transformers offer a compact, low-complexity solution for edge deployment with reduced bandwidth needs and acceptable accuracy. However, existing binary transformers perform inefficiently on current hardware due to the lack of binary specific optimizations. To address this, we introduce COBRA, an algorithm-architecture co-optimized binary Transformer accelerator for edge computing. COBRA features a real 1-bit binary multiplication unit, enabling matrix operations with -1, 0, and +1 values, surpassing ternary methods. With further hardware-friendly optimizations in the attention block, COBRA achieves up to 3,894.7 GOPS throughput and 448.7 GOPS/Watt energy efficiency on edge FPGAs, delivering a 311x energy efficiency improvement over GPUs and a 3.5x throughput improvement over the state-of-the-art binary accelerator, with only negligible inference accuracy degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_16269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
Qiao, Ye
Chen, Zhiheng
Wang, Yian
Zhang, Yifan
Deng, Yunzhe
Huang, Sitao
Hardware Architecture
Machine Learning
Transformer-based models have demonstrated superior performance in various fields, including natural language processing and computer vision. However, their enormous model size and high demands in computation, memory, and communication limit their deployment to edge platforms for local, secure inference. Binary transformers offer a compact, low-complexity solution for edge deployment with reduced bandwidth needs and acceptable accuracy. However, existing binary transformers perform inefficiently on current hardware due to the lack of binary specific optimizations. To address this, we introduce COBRA, an algorithm-architecture co-optimized binary Transformer accelerator for edge computing. COBRA features a real 1-bit binary multiplication unit, enabling matrix operations with -1, 0, and +1 values, surpassing ternary methods. With further hardware-friendly optimizations in the attention block, COBRA achieves up to 3,894.7 GOPS throughput and 448.7 GOPS/Watt energy efficiency on edge FPGAs, delivering a 311x energy efficiency improvement over GPUs and a 3.5x throughput improvement over the state-of-the-art binary accelerator, with only negligible inference accuracy degradation.
title COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
topic Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2504.16269