Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Uchino, Yuki, Ozaki, Katsuhisa, Imamura, Toshiyuki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910919819264000
author Uchino, Yuki
Ozaki, Katsuhisa
Imamura, Toshiyuki
author_facet Uchino, Yuki
Ozaki, Katsuhisa
Imamura, Toshiyuki
contents This study was aimed at simultaneously achieving sufficient accuracy and high performance for general matrix multiplications. Recent architectures, such as NVIDIA GPUs, feature high-performance units designed for low-precision matrix multiplications in machine learning models, and next-generation architectures are expected to follow the same design principle. The key to achieving superior performance is to fully leverage such architectures. The Ozaki scheme, a highly accurate matrix multiplication algorithm using error-free transformations, enables higher-precision matrix multiplication to be performed through multiple lower-precision matrix multiplications and higher-precision matrix additions. Ootomo et al. implemented the Ozaki scheme on high-performance matrix multiplication units with the aim of achieving both sufficient accuracy and high performance. This paper proposes alternative approaches to improving performance by reducing the numbers of lower-precision matrix multiplications and higher-precision matrix additions. Numerical experiments demonstrate the accuracy of the results and conduct performance benchmarks of the proposed approaches. These approaches are expected to yield more efficient results in next-generation architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit
Uchino, Yuki
Ozaki, Katsuhisa
Imamura, Toshiyuki
Distributed, Parallel, and Cluster Computing
This study was aimed at simultaneously achieving sufficient accuracy and high performance for general matrix multiplications. Recent architectures, such as NVIDIA GPUs, feature high-performance units designed for low-precision matrix multiplications in machine learning models, and next-generation architectures are expected to follow the same design principle. The key to achieving superior performance is to fully leverage such architectures. The Ozaki scheme, a highly accurate matrix multiplication algorithm using error-free transformations, enables higher-precision matrix multiplication to be performed through multiple lower-precision matrix multiplications and higher-precision matrix additions. Ootomo et al. implemented the Ozaki scheme on high-performance matrix multiplication units with the aim of achieving both sufficient accuracy and high performance. This paper proposes alternative approaches to improving performance by reducing the numbers of lower-precision matrix multiplications and higher-precision matrix additions. Numerical experiments demonstrate the accuracy of the results and conduct performance benchmarks of the proposed approaches. These approaches are expected to yield more efficient results in next-generation architectures.
title Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2409.13313