Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Haibin, Li, Wenming, Yan, Kai, Fan, Zhihua, Wu, Peiyang, Liu, Yuqun, Liu, Yanhuan, Qiang, Ziqing, Wu, Meng, Liu, Kunming, Ye, Xiaochun, Fan, Dongrui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913584356786176
author Wu, Haibin
Li, Wenming
Yan, Kai
Fan, Zhihua
Wu, Peiyang
Liu, Yuqun
Liu, Yanhuan
Qiang, Ziqing
Wu, Meng
Liu, Kunming
Ye, Xiaochun
Fan, Dongrui
author_facet Wu, Haibin
Li, Wenming
Yan, Kai
Fan, Zhihua
Wu, Peiyang
Liu, Yuqun
Liu, Yanhuan
Qiang, Ziqing
Wu, Meng
Liu, Kunming
Ye, Xiaochun
Fan, Dongrui
contents Recent neural networks (NNs) with self-attention exhibit competitiveness across different AI domains, but the essential attention mechanism brings massive computation and memory demands. To this end, various sparsity patterns are introduced to reduce the quadratic computation complexity, among which the structured butterfly sparsity has been proven efficient in computation reduction while maintaining model accuracy. However, its complicated data accessing pattern brings utilization degradation and makes parallelism hard to exploit in general block-oriented architecture like GPU. Since the reconfigurable dataflow architecture is known to have better data reusability and architectural flexibility in general NN-based acceleration, we want to apply it to the butterfly sparsity for acquiring better computational efficiency for attention workloads. We first propose a hybrid butterfly-sparsity network to obtain better trade-offs between attention accuracy and performance. Next, we propose a scalable multilayer dataflow method supported by coarse-grained streaming parallelism designs, to orchestrate the butterfly sparsity computation on the dataflow array. The experiments show that compared with Jetson Xavier NX, our design has a speedup of up to $14.34\times$ ($9.29\times$ on average) as well as $11.14\times$ energy efficiency advancement in attention workloads. In comparison with SOTA attention accelerators of the same peak performance, our dataflow architecture acquires $2.38\times$-$4.7\times$ efficiency improvement as well as $6.60\times$-$15.37\times$ energy reduction with butterfly sparsity optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00734
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
Wu, Haibin
Li, Wenming
Yan, Kai
Fan, Zhihua
Wu, Peiyang
Liu, Yuqun
Liu, Yanhuan
Qiang, Ziqing
Wu, Meng
Liu, Kunming
Ye, Xiaochun
Fan, Dongrui
Hardware Architecture
Recent neural networks (NNs) with self-attention exhibit competitiveness across different AI domains, but the essential attention mechanism brings massive computation and memory demands. To this end, various sparsity patterns are introduced to reduce the quadratic computation complexity, among which the structured butterfly sparsity has been proven efficient in computation reduction while maintaining model accuracy. However, its complicated data accessing pattern brings utilization degradation and makes parallelism hard to exploit in general block-oriented architecture like GPU. Since the reconfigurable dataflow architecture is known to have better data reusability and architectural flexibility in general NN-based acceleration, we want to apply it to the butterfly sparsity for acquiring better computational efficiency for attention workloads. We first propose a hybrid butterfly-sparsity network to obtain better trade-offs between attention accuracy and performance. Next, we propose a scalable multilayer dataflow method supported by coarse-grained streaming parallelism designs, to orchestrate the butterfly sparsity computation on the dataflow array. The experiments show that compared with Jetson Xavier NX, our design has a speedup of up to $14.34\times$ ($9.29\times$ on average) as well as $11.14\times$ energy efficiency advancement in attention workloads. In comparison with SOTA attention accelerators of the same peak performance, our dataflow architecture acquires $2.38\times$-$4.7\times$ efficiency improvement as well as $6.60\times$-$15.37\times$ energy reduction with butterfly sparsity optimization.
title Multilayer Dataflow: Orchestrate Butterfly Sparsity to Accelerate Attention Computation
topic Hardware Architecture
url https://arxiv.org/abs/2411.00734