Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Uchino, Yuki, Ozaki, Katsuhisa, Imamura, Toshiyuki
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911568315285504
author Uchino, Yuki
Ozaki, Katsuhisa
Imamura, Toshiyuki
author_facet Uchino, Yuki
Ozaki, Katsuhisa
Imamura, Toshiyuki
contents In this paper, we propose a method for emulating double-precision general matrix--matrix multiplication (DGEMM), a fundamental and performance-critical kernel in many high-performance computing applications. Ozaki-I and Ozaki-II are established DGEMM emulation schemes via low-precision matrix multiply-accumulate (MMA) units. For the Ozaki-I scheme, INT8-, FP8-, and FP16-based implementations have been proposed, all of which can be realized based on the same underlying algorithmic structure. In contrast, although INT8-based implementations of the Ozaki-II scheme have been reported, the original algorithm cannot be directly adapted to exploit FP8 MMA units. In several recent architectures, such as NVIDIA Blackwell Ultra and NVIDIA Rubin, INT8 performance has been reduced, making reliance on INT8 alone insufficient. Therefore, we introduce a novel technique to demonstrate DGEMM emulation based on the Ozaki-II scheme that operates on FP8 MMA units. Compared to the FP8-based Ozaki-I scheme, our method significantly reduces the computational cost and enables efficient FP64 emulation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10634
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
Uchino, Yuki
Ozaki, Katsuhisa
Imamura, Toshiyuki
Distributed, Parallel, and Cluster Computing
In this paper, we propose a method for emulating double-precision general matrix--matrix multiplication (DGEMM), a fundamental and performance-critical kernel in many high-performance computing applications. Ozaki-I and Ozaki-II are established DGEMM emulation schemes via low-precision matrix multiply-accumulate (MMA) units. For the Ozaki-I scheme, INT8-, FP8-, and FP16-based implementations have been proposed, all of which can be realized based on the same underlying algorithmic structure. In contrast, although INT8-based implementations of the Ozaki-II scheme have been reported, the original algorithm cannot be directly adapted to exploit FP8 MMA units. In several recent architectures, such as NVIDIA Blackwell Ultra and NVIDIA Rubin, INT8 performance has been reduced, making reliance on INT8 alone insufficient. Therefore, we introduce a novel technique to demonstrate DGEMM emulation based on the Ozaki-II scheme that operates on FP8 MMA units. Compared to the FP8-based Ozaki-I scheme, our method significantly reduces the computational cost and enables efficient FP64 emulation.
title Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2603.10634