Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Uchino, Yuki, Ozaki, Katsuhisa, Imamura, Toshiyuki |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Error Analysis of Matrix Multiplication Emulation Using Ozaki-II Scheme
by: Uchino, Yuki, et al.
Published: (2026)
by: Uchino, Yuki, et al.
Published: (2026)
Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit
by: Uchino, Yuki, et al.
Published: (2024)
by: Uchino, Yuki, et al.
Published: (2024)
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
by: Uchino, Yuki, et al.
Published: (2025)
by: Uchino, Yuki, et al.
Published: (2025)
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
by: Uchino, Yuki, et al.
Published: (2025)
by: Uchino, Yuki, et al.
Published: (2025)
DGEMM on Integer Matrix Multiplication Unit
by: Ootomo, Hiroyuki, et al.
Published: (2023)
by: Ootomo, Hiroyuki, et al.
Published: (2023)
Guaranteed DGEMM Accuracy While Using Reduced Precision Tensor Cores Through Extensions of the Ozaki Scheme
by: Schwarz, Angelika, et al.
Published: (2025)
by: Schwarz, Angelika, et al.
Published: (2025)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
by: Xue, Weicheng, et al.
Published: (2025)
by: Xue, Weicheng, et al.
Published: (2025)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
by: Lee, Haeun, et al.
Published: (2025)
by: Lee, Haeun, et al.
Published: (2025)
A Pilot Study on Tunable Precision Emulation via Automatic BLAS Offloading
by: Liu, Hang, et al.
Published: (2025)
by: Liu, Hang, et al.
Published: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
by: Jangda, Abhinav, et al.
Published: (2024)
by: Jangda, Abhinav, et al.
Published: (2024)
Ozaki Scheme II: A GEMM-oriented emulation of floating-point matrix multiplication using an integer modular technique
by: Ozaki, Katsuhisa, et al.
Published: (2025)
by: Ozaki, Katsuhisa, et al.
Published: (2025)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
by: Metere, Alfredo
Published: (2025)
by: Metere, Alfredo
Published: (2025)
On the Extension of Private Distributed Matrix Multiplication Schemes to the Grid Partition
by: Hofmeister, Christoph, et al.
Published: (2026)
by: Hofmeister, Christoph, et al.
Published: (2026)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
by: Qian, Matthew, et al.
Published: (2026)
by: Qian, Matthew, et al.
Published: (2026)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
by: Li, Shiju, et al.
Published: (2025)
by: Li, Shiju, et al.
Published: (2025)
W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs
by: He, Yuanhong, et al.
Published: (2026)
by: He, Yuanhong, et al.
Published: (2026)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
Execution-Centric Characterization of FP8 Matrix Cores, Asynchronous Execution, and Structured Sparsity on AMD MI300A
by: Jarmusch, Aaron, et al.
Published: (2026)
by: Jarmusch, Aaron, et al.
Published: (2026)
Improving Locality in Sparse and Dense Matrix Multiplications
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
by: Zhang, Lixing, et al.
Published: (2026)
by: Zhang, Lixing, et al.
Published: (2026)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
by: Remke, Stefan, et al.
Published: (2024)
by: Remke, Stefan, et al.
Published: (2024)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
by: Ranawaka, Isuru, et al.
Published: (2024)
by: Ranawaka, Isuru, et al.
Published: (2024)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
by: Adefemi, Temitayo
Published: (2024)
by: Adefemi, Temitayo
Published: (2024)
Demystifying ARM SME to Optimize General Matrix Multiplications
by: Deng, Chencheng, et al.
Published: (2025)
by: Deng, Chencheng, et al.
Published: (2025)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
by: Li, Zhonggen, et al.
Published: (2024)
by: Li, Zhonggen, et al.
Published: (2024)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
by: Li, Aiying, et al.
Published: (2026)
by: Li, Aiying, et al.
Published: (2026)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
by: Lei, Kelun, et al.
Published: (2025)
by: Lei, Kelun, et al.
Published: (2025)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026)
by: Shah, Milan, et al.
Published: (2026)
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling
by: Da, Wei, et al.
Published: (2026)
by: Da, Wei, et al.
Published: (2026)
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
by: Guo, Han, et al.
Published: (2024)
by: Guo, Han, et al.
Published: (2024)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
by: Park, Gunho, et al.
Published: (2022)
by: Park, Gunho, et al.
Published: (2022)
Efficiently Parallelizable Strassen-Based Multiplication of a Matrix by its Transpose
by: Arrigoni, Viviana, et al.
Published: (2021)
by: Arrigoni, Viviana, et al.
Published: (2021)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
by: Asudeh, Omid, et al.
Published: (2025)
by: Asudeh, Omid, et al.
Published: (2025)
Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
by: Zhang, Qiao, et al.
Published: (2025)
by: Zhang, Qiao, et al.
Published: (2025)
Matrix Multiplication in the MPC Model
by: Joshi, Lakshya, et al.
Published: (2025)
by: Joshi, Lakshya, et al.
Published: (2025)
BouquetFL: Emulating diverse participant hardware in Federated Learning
by: Geimer, Arno
Published: (2026)
by: Geimer, Arno
Published: (2026)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
by: Lei, Jie, et al.
Published: (2024)
by: Lei, Jie, et al.
Published: (2024)
Emulating a computing grid in a local environment for feature evaluation
by: Kalawana, Jananga, et al.
Published: (2024)
by: Kalawana, Jananga, et al.
Published: (2024)
Similar Items
-
Error Analysis of Matrix Multiplication Emulation Using Ozaki-II Scheme
by: Uchino, Yuki, et al.
Published: (2026) -
Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit
by: Uchino, Yuki, et al.
Published: (2024) -
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
by: Uchino, Yuki, et al.
Published: (2025) -
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
by: Uchino, Yuki, et al.
Published: (2025) -
DGEMM on Integer Matrix Multiplication Unit
by: Ootomo, Hiroyuki, et al.
Published: (2023)