PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Huizheng, Wang, Hongbin, Wang, Zichuan, Yue, Zhiheng, Wang, Yang, Li, Chao, Hu, Yang, Yin, Shouyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912813593657344
author Wang, Huizheng
Wang, Hongbin
Wang, Zichuan
Yue, Zhiheng
Wang, Yang
Li, Chao
Hu, Yang
Yin, Shouyi
author_facet Wang, Huizheng
Wang, Hongbin
Wang, Zichuan
Yue, Zhiheng
Wang, Yang
Li, Chao
Hu, Yang
Yin, Shouyi
contents Attention-based models have revolutionized AI, but the quadratic cost of self-attention incurs severe computational and memory overhead. Sparse attention methods alleviate this by skipping low-relevance token pairs. However, current approaches lack practicality due to the heavy expense of added sparsity predictor, which severely drops their hardware efficiency. This paper advances the state-of-the-art (SOTA) by proposing a bit-serial enable stage-fusion (BSF) mechanism, which eliminates the need for a separate predictor. However, it faces key challenges: 1) Inaccurate bit-sliced sparsity speculation leads to incorrect pruning; 2) Hardware under-utilization due to fine-grained and imbalanced bit-level workloads. 3) Tiling difficulty caused by the row-wise dependency in sparsity pruning criteria. We propose PADE, a predictor-free algorithm-hardware co-design for dynamic sparse attention acceleration. PADE features three key innovations: 1) Bit-wise uncertainty interval-enabled guard filtering (BUI-GF) strategy to accurately identify trivial tokens during each bit round; 2) Bidirectional sparsity-based out-of-order execution (BS-OOE) to improve hardware utilization; 3) Interleaving-based sparsity-tiled attention (ISTA) to reduce both I/O and computational complexity. These techniques, combined with custom accelerator designs, enable practical sparsity acceleration without relying on an added sparsity predictor. Extensive experiments on 22 benchmarks show that PADE achieves 7.43x speed up and 31.1x higher energy efficiency than Nvidia H100 GPU. Compared to SOTA accelerators, PADE achieves 5.1x, 4.3x and 3.4x energy saving than Sanger, DOTA and SOFA.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14322
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
Wang, Huizheng
Wang, Hongbin
Wang, Zichuan
Yue, Zhiheng
Wang, Yang
Li, Chao
Hu, Yang
Yin, Shouyi
Hardware Architecture
Signal Processing
Attention-based models have revolutionized AI, but the quadratic cost of self-attention incurs severe computational and memory overhead. Sparse attention methods alleviate this by skipping low-relevance token pairs. However, current approaches lack practicality due to the heavy expense of added sparsity predictor, which severely drops their hardware efficiency. This paper advances the state-of-the-art (SOTA) by proposing a bit-serial enable stage-fusion (BSF) mechanism, which eliminates the need for a separate predictor. However, it faces key challenges: 1) Inaccurate bit-sliced sparsity speculation leads to incorrect pruning; 2) Hardware under-utilization due to fine-grained and imbalanced bit-level workloads. 3) Tiling difficulty caused by the row-wise dependency in sparsity pruning criteria. We propose PADE, a predictor-free algorithm-hardware co-design for dynamic sparse attention acceleration. PADE features three key innovations: 1) Bit-wise uncertainty interval-enabled guard filtering (BUI-GF) strategy to accurately identify trivial tokens during each bit round; 2) Bidirectional sparsity-based out-of-order execution (BS-OOE) to improve hardware utilization; 3) Interleaving-based sparsity-tiled attention (ISTA) to reduce both I/O and computational complexity. These techniques, combined with custom accelerator designs, enable practical sparsity acceleration without relying on an added sparsity predictor. Extensive experiments on 22 benchmarks show that PADE achieves 7.43x speed up and 31.1x higher energy efficiency than Nvidia H100 GPU. Compared to SOTA accelerators, PADE achieves 5.1x, 4.3x and 3.4x energy saving than Sanger, DOTA and SOFA.
title PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage Fusion
topic Hardware Architecture
Signal Processing
url https://arxiv.org/abs/2512.14322