High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
Fuente:
arXiv
Saved in:
| Main Authors: | Uchino, Yuki, Ozaki, Katsuhisa, Imamura, Toshiyuki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
by: Uchino, Yuki, et al.
Published: (2026)
by: Uchino, Yuki, et al.
Published: (2026)
Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit
by: Uchino, Yuki, et al.
Published: (2024)
by: Uchino, Yuki, et al.
Published: (2024)
Error Analysis of Matrix Multiplication Emulation Using Ozaki-II Scheme
by: Uchino, Yuki, et al.
Published: (2026)
by: Uchino, Yuki, et al.
Published: (2026)
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
by: Uchino, Yuki, et al.
Published: (2025)
by: Uchino, Yuki, et al.
Published: (2025)
DGEMM on Integer Matrix Multiplication Unit
by: Ootomo, Hiroyuki, et al.
Published: (2023)
by: Ootomo, Hiroyuki, et al.
Published: (2023)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
by: Lei, Kelun, et al.
Published: (2025)
by: Lei, Kelun, et al.
Published: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
by: Jangda, Abhinav, et al.
Published: (2024)
by: Jangda, Abhinav, et al.
Published: (2024)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
by: Zhang, Lixing, et al.
Published: (2026)
by: Zhang, Lixing, et al.
Published: (2026)
Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
by: Adefemi, Temitayo
Published: (2024)
by: Adefemi, Temitayo
Published: (2024)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
by: Qian, Matthew, et al.
Published: (2026)
by: Qian, Matthew, et al.
Published: (2026)
Matrix-PIC: Harnessing Matrix Outer-product for High-Performance Particle-in-Cell Simulations
by: Rao, Yizhuo, et al.
Published: (2026)
by: Rao, Yizhuo, et al.
Published: (2026)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
by: Li, Shiju, et al.
Published: (2025)
by: Li, Shiju, et al.
Published: (2025)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
by: Wolfson-Pou, Jordi, et al.
Published: (2025)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
Efficiently Parallelizable Strassen-Based Multiplication of a Matrix by its Transpose
by: Arrigoni, Viviana, et al.
Published: (2021)
by: Arrigoni, Viviana, et al.
Published: (2021)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
by: Remke, Stefan, et al.
Published: (2024)
by: Remke, Stefan, et al.
Published: (2024)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
by: Ranawaka, Isuru, et al.
Published: (2024)
by: Ranawaka, Isuru, et al.
Published: (2024)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
by: Li, Zhonggen, et al.
Published: (2024)
by: Li, Zhonggen, et al.
Published: (2024)
Improving Locality in Sparse and Dense Matrix Multiplications
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
by: Asudeh, Omid, et al.
Published: (2025)
by: Asudeh, Omid, et al.
Published: (2025)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
by: Li, Aiying, et al.
Published: (2026)
by: Li, Aiying, et al.
Published: (2026)
Demystifying ARM SME to Optimize General Matrix Multiplications
by: Deng, Chencheng, et al.
Published: (2025)
by: Deng, Chencheng, et al.
Published: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026)
by: Shah, Milan, et al.
Published: (2026)
Matrix Multiplication in the MPC Model
by: Joshi, Lakshya, et al.
Published: (2025)
by: Joshi, Lakshya, et al.
Published: (2025)
MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit
by: Wang, Yinuo, et al.
Published: (2025)
by: Wang, Yinuo, et al.
Published: (2025)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
by: Zhang, Zhekai, et al.
Published: (2020)
by: Zhang, Zhekai, et al.
Published: (2020)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
by: Xue, Weicheng, et al.
Published: (2025)
by: Xue, Weicheng, et al.
Published: (2025)
Stencil Matrixization
by: Zhao, Wenxuan, et al.
Published: (2023)
by: Zhao, Wenxuan, et al.
Published: (2023)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
by: Li, Chendi, et al.
Published: (2022)
by: Li, Chendi, et al.
Published: (2022)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
by: Chang, Dali, et al.
Published: (2026)
by: Chang, Dali, et al.
Published: (2026)
Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
by: Gianinazzi, Lukas, et al.
Published: (2024)
by: Gianinazzi, Lukas, et al.
Published: (2024)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
Multivariate Polynomial Codes for Efficient Matrix Chain Multiplication in Distributed Systems
by: Gómez-Vilardebò, Jesús
Published: (2026)
by: Gómez-Vilardebò, Jesús
Published: (2026)
Mapping Parallel Matrix Multiplication in GotoBLAS2 to the AMD Versal ACAP for Deep Learning
by: Lei, Jie, et al.
Published: (2024)
by: Lei, Jie, et al.
Published: (2024)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
Tensor-Parallel Emulation of Quantum Circuits with Block-Cyclic Distributed Matrix Product States
by: Adamski, Jakub, et al.
Published: (2025)
by: Adamski, Jakub, et al.
Published: (2025)
Low-ordered Orthogonal Voxel Finite Element with INT8 Tensor Cores for GPU-based Explicit Elastic Wave Propagation Analysis
by: Ichimura, Tsuyoshi, et al.
Published: (2024)
by: Ichimura, Tsuyoshi, et al.
Published: (2024)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
by: Lacey, Dane C., et al.
Published: (2024)
by: Lacey, Dane C., et al.
Published: (2024)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
by: Zhuang, Chen, et al.
Published: (2025)
by: Zhuang, Chen, et al.
Published: (2025)
Similar Items
-
Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization
by: Uchino, Yuki, et al.
Published: (2026) -
Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit
by: Uchino, Yuki, et al.
Published: (2024) -
Error Analysis of Matrix Multiplication Emulation Using Ozaki-II Scheme
by: Uchino, Yuki, et al.
Published: (2026) -
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
by: Uchino, Yuki, et al.
Published: (2025) -
DGEMM on Integer Matrix Multiplication Unit
by: Ootomo, Hiroyuki, et al.
Published: (2023)