Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit
Fuente:
arXiv
Saved in:
| Main Authors: | Uchino, Yuki, Ozaki, Katsuhisa, Imamura, Toshiyuki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
by: Uchino, Yuki, et al.
Published: (2026)
by: Uchino, Yuki, et al.
Published: (2026)
Error Analysis of Matrix Multiplication Emulation Using Ozaki-II Scheme
by: Uchino, Yuki, et al.
Published: (2026)
by: Uchino, Yuki, et al.
Published: (2026)
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
by: Uchino, Yuki, et al.
Published: (2025)
by: Uchino, Yuki, et al.
Published: (2025)
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
by: Uchino, Yuki, et al.
Published: (2025)
by: Uchino, Yuki, et al.
Published: (2025)
DGEMM on Integer Matrix Multiplication Unit
by: Ootomo, Hiroyuki, et al.
Published: (2023)
by: Ootomo, Hiroyuki, et al.
Published: (2023)
Guaranteed DGEMM Accuracy While Using Reduced Precision Tensor Cores Through Extensions of the Ozaki Scheme
by: Schwarz, Angelika, et al.
Published: (2025)
by: Schwarz, Angelika, et al.
Published: (2025)
Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
by: Adefemi, Temitayo
Published: (2024)
by: Adefemi, Temitayo
Published: (2024)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
by: Lei, Kelun, et al.
Published: (2025)
by: Lei, Kelun, et al.
Published: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
by: Jangda, Abhinav, et al.
Published: (2024)
by: Jangda, Abhinav, et al.
Published: (2024)
On the Extension of Private Distributed Matrix Multiplication Schemes to the Grid Partition
by: Hofmeister, Christoph, et al.
Published: (2026)
by: Hofmeister, Christoph, et al.
Published: (2026)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
by: Qian, Matthew, et al.
Published: (2026)
by: Qian, Matthew, et al.
Published: (2026)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
by: Li, Shiju, et al.
Published: (2025)
by: Li, Shiju, et al.
Published: (2025)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
Improving Locality in Sparse and Dense Matrix Multiplications
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit
by: Wang, Yinuo, et al.
Published: (2025)
by: Wang, Yinuo, et al.
Published: (2025)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
by: Remke, Stefan, et al.
Published: (2024)
by: Remke, Stefan, et al.
Published: (2024)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
by: Ranawaka, Isuru, et al.
Published: (2024)
by: Ranawaka, Isuru, et al.
Published: (2024)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
by: Zhang, Lixing, et al.
Published: (2026)
by: Zhang, Lixing, et al.
Published: (2026)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
Demystifying ARM SME to Optimize General Matrix Multiplications
by: Deng, Chencheng, et al.
Published: (2025)
by: Deng, Chencheng, et al.
Published: (2025)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
by: Li, Zhonggen, et al.
Published: (2024)
by: Li, Zhonggen, et al.
Published: (2024)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
by: Li, Aiying, et al.
Published: (2026)
by: Li, Aiying, et al.
Published: (2026)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026)
by: Shah, Milan, et al.
Published: (2026)
Matrix-PIC: Harnessing Matrix Outer-product for High-Performance Particle-in-Cell Simulations
by: Rao, Yizhuo, et al.
Published: (2026)
by: Rao, Yizhuo, et al.
Published: (2026)
Ozaki Scheme II: A GEMM-oriented emulation of floating-point matrix multiplication using an integer modular technique
by: Ozaki, Katsuhisa, et al.
Published: (2025)
by: Ozaki, Katsuhisa, et al.
Published: (2025)
Efficiently Parallelizable Strassen-Based Multiplication of a Matrix by its Transpose
by: Arrigoni, Viviana, et al.
Published: (2021)
by: Arrigoni, Viviana, et al.
Published: (2021)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
by: Asudeh, Omid, et al.
Published: (2025)
by: Asudeh, Omid, et al.
Published: (2025)
Matrix Multiplication in the MPC Model
by: Joshi, Lakshya, et al.
Published: (2025)
by: Joshi, Lakshya, et al.
Published: (2025)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
by: Lei, Jie, et al.
Published: (2024)
by: Lei, Jie, et al.
Published: (2024)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
Overview and Prospects of Using Integer Surrogate Keys for Data Warehouse Performance Optimization
by: Stumpf, Sviatoslav, et al.
Published: (2025)
by: Stumpf, Sviatoslav, et al.
Published: (2025)
PARS3: Parallel Sparse Skew-Symmetric Matrix-Vector Multiplication with Reverse Cuthill-McKee Reordering
by: Yildirim, Selin, et al.
Published: (2024)
by: Yildirim, Selin, et al.
Published: (2024)
Maxing Out the SVM: Performance Impact of Memory and Program Cache Sizes in the Agave Validator
by: Vural, Turan, et al.
Published: (2025)
by: Vural, Turan, et al.
Published: (2025)
W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs
by: He, Yuanhong, et al.
Published: (2026)
by: He, Yuanhong, et al.
Published: (2026)
scaleTRIM: Scalable TRuncation-Based Integer Approximate Multiplier with Linearization and Compensation
by: Farahmand, Ebrahim, et al.
Published: (2023)
by: Farahmand, Ebrahim, et al.
Published: (2023)
Multithreaded Fine-Grained Asynchronous BSP for Integer Sorting with LCI and OpenMP
by: Cheng, Minyu, et al.
Published: (2026)
by: Cheng, Minyu, et al.
Published: (2026)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
by: Zhuang, Chen, et al.
Published: (2025)
by: Zhuang, Chen, et al.
Published: (2025)
Stencil Matrixization
by: Zhao, Wenxuan, et al.
Published: (2023)
by: Zhao, Wenxuan, et al.
Published: (2023)
Performance analysis of mdx II: A next-generation cloud platform for cross-disciplinary data science research
by: Takahashi, Keichi, et al.
Published: (2025)
by: Takahashi, Keichi, et al.
Published: (2025)
Similar Items
-
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
by: Uchino, Yuki, et al.
Published: (2026) -
Error Analysis of Matrix Multiplication Emulation Using Ozaki-II Scheme
by: Uchino, Yuki, et al.
Published: (2026) -
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
by: Uchino, Yuki, et al.
Published: (2025) -
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
by: Uchino, Yuki, et al.
Published: (2025) -
DGEMM on Integer Matrix Multiplication Unit
by: Ootomo, Hiroyuki, et al.
Published: (2023)