BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Zisheng, He, Xiaoyu, Song, Maoyuan, Qiu, Guoliang, Liao, Chao, Wu, Chen, Sun, Yonggang, Li, Zhichun, Xie, Xiaoru, Luo, Yuanyong, Liu, Hu, Lu, Pinyan, Liao, Heng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910008768200704
author Ye, Zisheng
He, Xiaoyu
Song, Maoyuan
Qiu, Guoliang
Liao, Chao
Wu, Chen
Sun, Yonggang
Li, Zhichun
Xie, Xiaoru
Luo, Yuanyong
Liu, Hu
Lu, Pinyan
Liao, Heng
author_facet Ye, Zisheng
He, Xiaoyu
Song, Maoyuan
Qiu, Guoliang
Liao, Chao
Wu, Chen
Sun, Yonggang
Li, Zhichun
Xie, Xiaoru
Luo, Yuanyong
Liu, Hu
Lu, Pinyan
Liao, Heng
contents As the performance gains from accelerating quantized matrix multiplication plateau, the softmax operation becomes the critical bottleneck in Transformer inference. This bottleneck stems from two hardware limitations: (1) limited data bandwidth between matrix and vector compute cores, and (2) the significant area cost of high-precision (FP32/16) exponentiation units (EXP2). To address these issues, we introduce a novel low-precision workflow that employs a specific 8-bit floating-point format (HiF8) and block-aware precision rescaling for softmax. Crucially, our algorithmic innovations make low-precision softmax feasible without the significant model accuracy loss that hampers direct low-precision approaches. Specifically, our design (i) halves the required data movement bandwidth by enabling matrix multiplication outputs constrained to 8-bit, and (ii) substantially reduces the EXP2 unit area by computing exponentiations in low (8-bit) precision. Extensive evaluation on language models and multi-modal models confirms the validity of our method. By alleviating the vector computation bottleneck, our work paves the way for doubling end-to-end inference throughput without increasing chip area, and offers a concrete co-design path for future low-precision hardware and software.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02071
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling
Ye, Zisheng
He, Xiaoyu
Song, Maoyuan
Qiu, Guoliang
Liao, Chao
Wu, Chen
Sun, Yonggang
Li, Zhichun
Xie, Xiaoru
Luo, Yuanyong
Liu, Hu
Lu, Pinyan
Liao, Heng
Machine Learning
As the performance gains from accelerating quantized matrix multiplication plateau, the softmax operation becomes the critical bottleneck in Transformer inference. This bottleneck stems from two hardware limitations: (1) limited data bandwidth between matrix and vector compute cores, and (2) the significant area cost of high-precision (FP32/16) exponentiation units (EXP2). To address these issues, we introduce a novel low-precision workflow that employs a specific 8-bit floating-point format (HiF8) and block-aware precision rescaling for softmax. Crucially, our algorithmic innovations make low-precision softmax feasible without the significant model accuracy loss that hampers direct low-precision approaches. Specifically, our design (i) halves the required data movement bandwidth by enabling matrix multiplication outputs constrained to 8-bit, and (ii) substantially reduces the EXP2 unit area by computing exponentiations in low (8-bit) precision. Extensive evaluation on language models and multi-modal models confirms the validity of our method. By alleviating the vector computation bottleneck, our work paves the way for doubling end-to-end inference throughput without increasing chip area, and offers a concrete co-design path for future low-precision hardware and software.
title BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling
topic Machine Learning
url https://arxiv.org/abs/2602.02071