DGEMM on Integer Matrix Multiplication Unit

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ootomo, Hiroyuki, Ozaki, Katsuhisa, Yokota, Rio
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916184843091968
author Ootomo, Hiroyuki
Ozaki, Katsuhisa
Yokota, Rio
author_facet Ootomo, Hiroyuki
Ozaki, Katsuhisa
Yokota, Rio
contents Deep learning hardware achieves high throughput and low power consumption by reducing computing precision and specializing in matrix multiplication. For machine learning inference, fixed-point value computation is commonplace, where the input and output values and the model parameters are quantized. Thus, many processors are now equipped with fast integer matrix multiplication units (IMMU). It is of significant interest to find a way to harness these IMMUs to improve the performance of HPC applications while maintaining accuracy. We focus on the Ozaki scheme, which computes a high-precision matrix multiplication by using lower-precision computing units, and show the advantages and disadvantages of using IMMU. The experiment using integer Tensor Cores shows that we can compute double-precision matrix multiplication faster than cuBLAS and an existing Ozaki scheme implementation on FP16 Tensor Cores on NVIDIA consumer GPUs. Furthermore, we demonstrate accelerating a quantum circuit simulation by up to 4.33 while maintaining the FP64 accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2306_11975
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DGEMM on Integer Matrix Multiplication Unit
Ootomo, Hiroyuki
Ozaki, Katsuhisa
Yokota, Rio
Distributed, Parallel, and Cluster Computing
Deep learning hardware achieves high throughput and low power consumption by reducing computing precision and specializing in matrix multiplication. For machine learning inference, fixed-point value computation is commonplace, where the input and output values and the model parameters are quantized. Thus, many processors are now equipped with fast integer matrix multiplication units (IMMU). It is of significant interest to find a way to harness these IMMUs to improve the performance of HPC applications while maintaining accuracy. We focus on the Ozaki scheme, which computes a high-precision matrix multiplication by using lower-precision computing units, and show the advantages and disadvantages of using IMMU. The experiment using integer Tensor Cores shows that we can compute double-precision matrix multiplication faster than cuBLAS and an existing Ozaki scheme implementation on FP16 Tensor Cores on NVIDIA consumer GPUs. Furthermore, we demonstrate accelerating a quantum circuit simulation by up to 4.33 while maintaining the FP64 accuracy.
title DGEMM on Integer Matrix Multiplication Unit
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2306.11975