Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhao, Liang, Shao, Kunming, Tian, Fengshi, Cheng, Tim Kwang-Ting, Tsui, Chi-Ying, Zou, Yi
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2502.00687
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916593688117248
author Zhao, Liang
Shao, Kunming
Tian, Fengshi
Cheng, Tim Kwang-Ting
Tsui, Chi-Ying
Zou, Yi
author_facet Zhao, Liang
Shao, Kunming
Tian, Fengshi
Cheng, Tim Kwang-Ting
Tsui, Chi-Ying
Zou, Yi
contents Deploying mixed-precision neural networks on edge devices is friendly to hardware resources and power consumption. To support fully mixed-precision neural network inference, it is necessary to design flexible hardware accelerators for continuous varying precision operations. However, the previous works have issues on hardware utilization and overhead of reconfigurable logic. In this paper, we propose an efficient accelerator for 2~8-bit precision scaling with serial activation input and parallel weight preloaded. First, we set two loading modes for the weight operands and decompose the weight into the corresponding bitwidths, which extends the weight precision support efficiently. Then, to improve hardware utilization of low-precision operations, we design the architecture that performs bit-serial MAC operation with systolic dataflow, and the partial sums are combined spatially. Furthermore, we designed an efficient carry save adder tree supporting both signed and unsigned number summation across rows. The experiment result shows that the proposed accelerator, synthesized with TSMC 28nm CMOS technology, achieves peak throughput of 4.09TOPS and peak energy efficiency of 68.94TOPS/W at 2/2-bit operations.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight Combination
Zhao, Liang
Shao, Kunming
Tian, Fengshi
Cheng, Tim Kwang-Ting
Tsui, Chi-Ying
Zou, Yi
Hardware Architecture
Systems and Control
Deploying mixed-precision neural networks on edge devices is friendly to hardware resources and power consumption. To support fully mixed-precision neural network inference, it is necessary to design flexible hardware accelerators for continuous varying precision operations. However, the previous works have issues on hardware utilization and overhead of reconfigurable logic. In this paper, we propose an efficient accelerator for 2~8-bit precision scaling with serial activation input and parallel weight preloaded. First, we set two loading modes for the weight operands and decompose the weight into the corresponding bitwidths, which extends the weight precision support efficiently. Then, to improve hardware utilization of low-precision operations, we design the architecture that performs bit-serial MAC operation with systolic dataflow, and the partial sums are combined spatially. Furthermore, we designed an efficient carry save adder tree supporting both signed and unsigned number summation across rows. The experiment result shows that the proposed accelerator, synthesized with TSMC 28nm CMOS technology, achieves peak throughput of 4.09TOPS and peak energy efficiency of 68.94TOPS/W at 2/2-bit operations.
title A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight Combination
topic Hardware Architecture
Systems and Control
url https://arxiv.org/abs/2502.00687