FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kulp, Gabriel, Ensinger, Andrew, Chen, Lizhong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929327204990976
author Kulp, Gabriel
Ensinger, Andrew
Chen, Lizhong
author_facet Kulp, Gabriel
Ensinger, Andrew
Chen, Lizhong
contents Tensors play a vital role in machine learning (ML) and often exhibit properties best explored while maintaining high-order. Efficiently performing ML computations requires taking advantage of sparsity, but generalized hardware support is challenging. This paper introduces FLAASH, a flexible and modular accelerator design for sparse tensor contraction that achieves over 25x speedup for a deep learning workload. Our architecture performs sparse high-order tensor contraction by distributing sparse dot products, or portions thereof, to numerous Sparse Dot Product Engines (SDPEs). Memory structure and job distribution can be customized, and we demonstrate a simple approach as a proof of concept. We address the challenges associated with control flow to navigate data structures, high-order representation, and high-sparsity handling. The effectiveness of our approach is demonstrated through various evaluations, showcasing significant speedup as sparsity and order increase.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16317
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction
Kulp, Gabriel
Ensinger, Andrew
Chen, Lizhong
Hardware Architecture
Machine Learning
Tensors play a vital role in machine learning (ML) and often exhibit properties best explored while maintaining high-order. Efficiently performing ML computations requires taking advantage of sparsity, but generalized hardware support is challenging. This paper introduces FLAASH, a flexible and modular accelerator design for sparse tensor contraction that achieves over 25x speedup for a deep learning workload. Our architecture performs sparse high-order tensor contraction by distributing sparse dot products, or portions thereof, to numerous Sparse Dot Product Engines (SDPEs). Memory structure and job distribution can be customized, and we demonstrate a simple approach as a proof of concept. We address the challenges associated with control flow to navigate data structures, high-order representation, and high-sparsity handling. The effectiveness of our approach is demonstrated through various evaluations, showcasing significant speedup as sparsity and order increase.
title FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction
topic Hardware Architecture
Machine Learning
url https://arxiv.org/abs/2404.16317