APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909602909519872 |
|---|---|
| author | Tan, Yonghao Dong, Pingcheng Wu, Yongkun Liu, Yu Liu, Xuejiao Luo, Peng Liu, Shih-Yang Huang, Xijie Zhang, Dong Liang, Luhong Cheng, Kwang-Ting |
| author_facet | Tan, Yonghao Dong, Pingcheng Wu, Yongkun Liu, Yu Liu, Xuejiao Luo, Peng Liu, Shih-Yang Huang, Xijie Zhang, Dong Liang, Luhong Cheng, Kwang-Ting |
| contents | DNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of high-precision partial sums (PSUMs) leads to excessive memory demands in architectures utilizing input/weight stationary dataflows. Traditional compression strategies have typically overlooked PSUM quantization, which may account for 69% of power consumption. This study introduces a novel Additive Partial Sum Quantization (APSQ) method, seamlessly integrating PSUM accumulation into the quantization framework. A grouping strategy that combines APSQ with PSUM quantization enhanced by a reconfigurable architecture is further proposed. The APSQ performs nearly lossless on NLP and CV tasks across BERT, Segformer, and EfficientViT models while compressing PSUMs to INT8. This leads to a notable reduction in energy costs by 28-87%. Extended experiments on LLaMA2-7B demonstrate the potential of APSQ for large language models. Code is available at https://github.com/Yonghao-Tan/APSQ. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_03748 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design Tan, Yonghao Dong, Pingcheng Wu, Yongkun Liu, Yu Liu, Xuejiao Luo, Peng Liu, Shih-Yang Huang, Xijie Zhang, Dong Liang, Luhong Cheng, Kwang-Ting Hardware Architecture Artificial Intelligence DNN accelerators, significantly advanced by model compression and specialized dataflow techniques, have marked considerable progress. However, the frequent access of high-precision partial sums (PSUMs) leads to excessive memory demands in architectures utilizing input/weight stationary dataflows. Traditional compression strategies have typically overlooked PSUM quantization, which may account for 69% of power consumption. This study introduces a novel Additive Partial Sum Quantization (APSQ) method, seamlessly integrating PSUM accumulation into the quantization framework. A grouping strategy that combines APSQ with PSUM quantization enhanced by a reconfigurable architecture is further proposed. The APSQ performs nearly lossless on NLP and CV tasks across BERT, Segformer, and EfficientViT models while compressing PSUMs to INT8. This leads to a notable reduction in energy costs by 28-87%. Extended experiments on LLaMA2-7B demonstrate the potential of APSQ for large language models. Code is available at https://github.com/Yonghao-Tan/APSQ. |
| title | APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design |
| topic | Hardware Architecture Artificial Intelligence |
| url | https://arxiv.org/abs/2505.03748 |