cuTeSpMM: Accelerating Sparse-Dense Matrix Multiplication using GPU Tensor Cores
Fuente:
arXiv
Saved in:
| Main Authors: | Xiang, Lizhi, Asudeh, Omid, Sabin, Gerald, Sukumaran-Rajam, Aravind, Sadayappan, P. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
by: Asudeh, Omid, et al.
Published: (2025)
by: Asudeh, Omid, et al.
Published: (2025)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication
by: Dehghankar, Mohsen, et al.
Published: (2026)
by: Dehghankar, Mohsen, et al.
Published: (2026)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
by: Li, Zhonggen, et al.
Published: (2024)
by: Li, Zhonggen, et al.
Published: (2024)
CoNST: Code Generator for Sparse Tensor Networks
by: Raje, Saurabh, et al.
Published: (2024)
by: Raje, Saurabh, et al.
Published: (2024)
AutoSAGE: Input-Aware CUDA Scheduling for Sparse GNN Aggregation (SpMM/SDDMM) and CSR Attention
by: Stankovic, Aleksandar
Published: (2025)
by: Stankovic, Aleksandar
Published: (2025)
Libra: Unleashing GPU Heterogeneity for High-Performance Sparse Matrix Multiplication
by: Shi, Jinliang, et al.
Published: (2025)
by: Shi, Jinliang, et al.
Published: (2025)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
by: Anik, Md Saidul Hoque, et al.
Published: (2024)
by: Anik, Md Saidul Hoque, et al.
Published: (2024)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
by: Curless, Brian, et al.
Published: (2025)
by: Curless, Brian, et al.
Published: (2025)
High Performance Matrix Multiplication
by: Davis, Ethan
Published: (2025)
by: Davis, Ethan
Published: (2025)
Acceleration of Tensor-Product Operations with Tensor Cores
by: Cui, Cu
Published: (2024)
by: Cui, Cu
Published: (2024)
Benchmark-based Study of CPU/GPU Power-Related Features through JAX and TensorFlow
by: Tchakoute, Roblex Nana, et al.
Published: (2025)
by: Tchakoute, Roblex Nana, et al.
Published: (2025)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
by: Zhuang, Chen, et al.
Published: (2025)
by: Zhuang, Chen, et al.
Published: (2025)
Diagonally-Addressed Matrix Nicknack: How to improve SpMV performance
by: Saak, Jens, et al.
Published: (2023)
by: Saak, Jens, et al.
Published: (2023)
Accelerating Sparse Tensor Decomposition Using Adaptive Linearized Representation
by: Laukemann, Jan, et al.
Published: (2024)
by: Laukemann, Jan, et al.
Published: (2024)
Insum: Sparse GPU Kernels Simplified and Optimized with Indirect Einsums
by: Won, Jaeyeon, et al.
Published: (2025)
by: Won, Jaeyeon, et al.
Published: (2025)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
GPU Acceleration of Sparse Fully Homomorphic Encrypted DNNs
by: D'Agata, Lara, et al.
Published: (2026)
by: D'Agata, Lara, et al.
Published: (2026)
GPU-Accelerated Parallel Selected Inversion for Structured Matrices Using sTiles
by: Fattah, Esmail Abdul, et al.
Published: (2025)
by: Fattah, Esmail Abdul, et al.
Published: (2025)
FlexiSAGA: A Flexible Systolic Array GEMM Accelerator for Sparse and Dense Processing
by: Müller, Mika Markus, et al.
Published: (2025)
by: Müller, Mika Markus, et al.
Published: (2025)
Towards a Higher Roofline for Matrix-Vector Multiplication in Matrix-Free HOSFEM
by: Cao, Zijian, et al.
Published: (2025)
by: Cao, Zijian, et al.
Published: (2025)
In-Situ Techniques on GPU-Accelerated Data-Intensive Applications
by: Ju, Yi, et al.
Published: (2024)
by: Ju, Yi, et al.
Published: (2024)
GPU-Accelerated INT8 Quantization for KV Cache Compression in Large Language Models
by: Taneja, Maanas, et al.
Published: (2026)
by: Taneja, Maanas, et al.
Published: (2026)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
by: Hossain, Md Arafat, et al.
Published: (2025)
by: Hossain, Md Arafat, et al.
Published: (2025)
Accelerating High-Order Finite Element Simulations at Extreme Scale with FP64 Tensor Cores
by: Tu, Jiqun, et al.
Published: (2026)
by: Tu, Jiqun, et al.
Published: (2026)
Accelerating Sparse Ternary GEMM for Quantized ML on Apple Silicon
by: Lipshitz, Baraq, et al.
Published: (2025)
by: Lipshitz, Baraq, et al.
Published: (2025)
A Test for FLOPs as a Discriminant for Linear Algebra Algorithms
by: Sankaran, Aravind, et al.
Published: (2022)
by: Sankaran, Aravind, et al.
Published: (2022)
Inspection of I/O Operations from System Call Traces using Directly-Follows-Graph
by: Sankaran, Aravind, et al.
Published: (2024)
by: Sankaran, Aravind, et al.
Published: (2024)
Efficient Hardware Accelerator Based on Medium Granularity Dataflow for SpTRSV
by: Chen, Qian, et al.
Published: (2024)
by: Chen, Qian, et al.
Published: (2024)
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
by: Yao, Feiyu, et al.
Published: (2026)
by: Yao, Feiyu, et al.
Published: (2026)
sTiles: An Accelerated Computational Framework for Sparse Factorizations of Structured Matrices
by: Fattah, Esmail Abdul, et al.
Published: (2025)
by: Fattah, Esmail Abdul, et al.
Published: (2025)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
by: Zhang, Lingqi, et al.
Published: (2025)
by: Zhang, Lingqi, et al.
Published: (2025)
AMD MI300X GPU Performance Analysis
by: Ambati, Chandrish, et al.
Published: (2025)
by: Ambati, Chandrish, et al.
Published: (2025)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
by: Zhang, Lixing, et al.
Published: (2026)
by: Zhang, Lixing, et al.
Published: (2026)
GoldbachGPU: An Open Source GPU-Accelerated Framework for Verification of Goldbach's Conjecture
by: Llorente-Saguer, Isaac
Published: (2026)
by: Llorente-Saguer, Isaac
Published: (2026)
Ranking with Ties based on Noisy Performance Data
by: Sankaran, Aravind, et al.
Published: (2024)
by: Sankaran, Aravind, et al.
Published: (2024)
An Analytical Cost Model for Fast Evaluation of Multiple Compute-Engine CNN Accelerators
by: Qararyah, Fareed, et al.
Published: (2025)
by: Qararyah, Fareed, et al.
Published: (2025)
DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static Information
by: Wang, Qiang, et al.
Published: (2024)
by: Wang, Qiang, et al.
Published: (2024)
Karatsuba Matrix Multiplication and its Efficient Custom Hardware Implementations
by: Pogue, Trevor E., et al.
Published: (2025)
by: Pogue, Trevor E., et al.
Published: (2025)
AGIPC: Adaptive In-Solve Algebraic Coarsening for GPU IPC
by: Wang, Xuan, et al.
Published: (2026)
by: Wang, Xuan, et al.
Published: (2026)
Similar Items
-
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
by: Asudeh, Omid, et al.
Published: (2025) -
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025) -
RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication
by: Dehghankar, Mohsen, et al.
Published: (2026) -
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
by: Li, Zhonggen, et al.
Published: (2024) -
CoNST: Code Generator for Sparse Tensor Networks
by: Raje, Saurabh, et al.
Published: (2024)