HadaCore: Tensor Core Accelerated Hadamard Transform Kernel

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agarwal, Krish, Astra, Rishi, Hoque, Adnan, Srivatsa, Mudhakar, Ganti, Raghu, Wright, Less, Chen, Sijia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917866951933952
author Agarwal, Krish
Astra, Rishi
Hoque, Adnan
Srivatsa, Mudhakar
Ganti, Raghu
Wright, Less
Chen, Sijia
author_facet Agarwal, Krish
Astra, Rishi
Hoque, Adnan
Srivatsa, Mudhakar
Ganti, Raghu
Wright, Less
Chen, Sijia
contents We present HadaCore, a modified Fast Walsh-Hadamard Transform (FWHT) algorithm optimized for the Tensor Cores present in modern GPU hardware. HadaCore follows the recursive structure of the original FWHT algorithm, achieving the same asymptotic runtime complexity but leveraging a hardware-aware work decomposition that benefits from Tensor Core acceleration. This reduces bottlenecks from compute and data exchange. On Nvidia A100 and H100 GPUs, HadaCore achieves speedups of 1.1-1.4x and 1.0-1.3x, with a peak gain of 3.5x and 3.6x respectively, when compared to the existing state-of-the-art implementation of the original algorithm. We also show that when using FP16 or BF16, our implementation is numerically accurate, enabling comparable accuracy on MMLU benchmarks when used in an end-to-end Llama3 inference run with quantized (FP8) attention.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08832
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
Agarwal, Krish
Astra, Rishi
Hoque, Adnan
Srivatsa, Mudhakar
Ganti, Raghu
Wright, Less
Chen, Sijia
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
We present HadaCore, a modified Fast Walsh-Hadamard Transform (FWHT) algorithm optimized for the Tensor Cores present in modern GPU hardware. HadaCore follows the recursive structure of the original FWHT algorithm, achieving the same asymptotic runtime complexity but leveraging a hardware-aware work decomposition that benefits from Tensor Core acceleration. This reduces bottlenecks from compute and data exchange. On Nvidia A100 and H100 GPUs, HadaCore achieves speedups of 1.1-1.4x and 1.0-1.3x, with a peak gain of 3.5x and 3.6x respectively, when compared to the existing state-of-the-art implementation of the original algorithm. We also show that when using FP16 or BF16, our implementation is numerically accurate, enabling comparable accuracy on MMLU benchmarks when used in an end-to-end Llama3 inference run with quantized (FP8) attention.
title HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2412.08832