APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tan, Yonghao, Dong, Pingcheng, Wu, Yongkun, Liu, Yu, Liu, Xuejiao, Luo, Peng, Liu, Shih-Yang, Huang, Xijie, Zhang, Dong, Liang, Luhong, Cheng, Kwang-Ting
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909602909519872
author Tan, Yonghao
Dong, Pingcheng
Wu, Yongkun
Liu, Yu
Liu, Xuejiao
Luo, Peng
Liu, Shih-Yang
Huang, Xijie
Zhang, Dong
Liang, Luhong
Cheng, Kwang-Ting
author_facet Tan, Yonghao
Dong, Pingcheng
Wu, Yongkun
Liu, Yu
Liu, Xuejiao
Luo, Peng
Liu, Shih-Yang
Huang, Xijie
Zhang, Dong
Liang, Luhong
Cheng, Kwang-Ting
contents DNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of high-precision partial sums (PSUMs) leads to excessive memory demands in architectures utilizing input/weight stationary dataflows. Traditional compression strategies have typically overlooked PSUM quantization, which may account for 69% of power consumption. This study introduces a novel Additive Partial Sum Quantization (APSQ) method, seamlessly integrating PSUM accumulation into the quantization framework. A grouping strategy that combines APSQ with PSUM quantization enhanced by a reconfigurable architecture is further proposed. The APSQ performs nearly lossless on NLP and CV tasks across BERT, Segformer, and EfficientViT models while compressing PSUMs to INT8. This leads to a notable reduction in energy costs by 28-87%. Extended experiments on LLaMA2-7B demonstrate the potential of APSQ for large language models. Code is available at https://github.com/Yonghao-Tan/APSQ.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03748
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
Tan, Yonghao
Dong, Pingcheng
Wu, Yongkun
Liu, Yu
Liu, Xuejiao
Luo, Peng
Liu, Shih-Yang
Huang, Xijie
Zhang, Dong
Liang, Luhong
Cheng, Kwang-Ting
Hardware Architecture
Artificial Intelligence
DNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of high-precision partial sums (PSUMs) leads to excessive memory demands in architectures utilizing input/weight stationary dataflows. Traditional compression strategies have typically overlooked PSUM quantization, which may account for 69% of power consumption. This study introduces a novel Additive Partial Sum Quantization (APSQ) method, seamlessly integrating PSUM accumulation into the quantization framework. A grouping strategy that combines APSQ with PSUM quantization enhanced by a reconfigurable architecture is further proposed. The APSQ performs nearly lossless on NLP and CV tasks across BERT, Segformer, and EfficientViT models while compressing PSUMs to INT8. This leads to a notable reduction in energy costs by 28-87%. Extended experiments on LLaMA2-7B demonstrate the potential of APSQ for large language models. Code is available at https://github.com/Yonghao-Tan/APSQ.
title APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
topic Hardware Architecture
Artificial Intelligence
url https://arxiv.org/abs/2505.03748