Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Haisha, Li, San, Wang, Jiaheng, Zhou, Chunbao, Wang, Jue, Xin, Zhikuang, Li, Shunde, Liang, Zhiqiang, Pan, Zhijie, Liu, Fang, Zeng, Yan, Wang, Yangang, Chi, Xuebin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)
von: Li, Aiying, et al.
Veröffentlicht: (2026)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
von: Ma, Cong, et al.
Veröffentlicht: (2025)
von: Ma, Cong, et al.
Veröffentlicht: (2025)
AES-SpMM: Balancing Accuracy and Speed by Adaptive Edge Sampling Strategy to Accelerate SpMM in GNNs
von: Song, Yingchen, et al.
Veröffentlicht: (2025)
von: Song, Yingchen, et al.
Veröffentlicht: (2025)
Hybrid SpMM SC26 AD
von: Wang, Haoyu
Veröffentlicht: (2026)
von: Wang, Haoyu
Veröffentlicht: (2026)
cuTeSpMM: Accelerating Sparse-Dense Matrix Multiplication using GPU Tensor Cores
von: Xiang, Lizhi, et al.
Veröffentlicht: (2025)
von: Xiang, Lizhi, et al.
Veröffentlicht: (2025)
High Performance Unstructured SpMM Computation Using Tensor Cores
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
von: Suzuki, Kengo, et al.
Veröffentlicht: (2026)
von: Suzuki, Kengo, et al.
Veröffentlicht: (2026)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
GPU-Accelerated Algorithms for Process Mapping
von: Samoldekin, Petr, et al.
Veröffentlicht: (2025)
von: Samoldekin, Petr, et al.
Veröffentlicht: (2025)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
von: Kriemann, Ronald
Veröffentlicht: (2024)
von: Kriemann, Ronald
Veröffentlicht: (2024)
Characterization of $n$-Lie Derivations on Generalized Matrix Algebras
von: Liang, Xinfeng, et al.
Veröffentlicht: (2026)
von: Liang, Xinfeng, et al.
Veröffentlicht: (2026)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
Algorithms for Parallel Shared-Memory Sparse Matrix-Vector Multiplication on Unstructured Matrices
von: Bergmans, Kobe, et al.
Veröffentlicht: (2025)
von: Bergmans, Kobe, et al.
Veröffentlicht: (2025)
Robust Tensor CUR Decompositions: Rapid Low-Tucker-Rank Tensor Recovery with Sparse Corruption
von: Cai, HanQin, et al.
Veröffentlicht: (2023)
von: Cai, HanQin, et al.
Veröffentlicht: (2023)
AutoSAGE: Input-Aware CUDA Scheduling for Sparse GNN Aggregation (SpMM/SDDMM) and CSR Attention
von: Stankovic, Aleksandar
Veröffentlicht: (2025)
von: Stankovic, Aleksandar
Veröffentlicht: (2025)
Matrix equivalence to Smith normal form: new theoretical results for multivariate polynomial matrices
von: Lu, Dong, et al.
Veröffentlicht: (2026)
von: Lu, Dong, et al.
Veröffentlicht: (2026)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
von: Li, Bohan, et al.
Veröffentlicht: (2026)
von: Li, Bohan, et al.
Veröffentlicht: (2026)
A Sparse Tensor Generator with Efficient Feature Extraction
von: Torun, Tugba, et al.
Veröffentlicht: (2024)
von: Torun, Tugba, et al.
Veröffentlicht: (2024)
Matrix perturbation bounds via contour bootstrapping
von: Tran, Phuc, et al.
Veröffentlicht: (2024)
von: Tran, Phuc, et al.
Veröffentlicht: (2024)
Compilation of Generalized Matrix Chains with Symbolic Sizes
von: López, Francisco, et al.
Veröffentlicht: (2025)
von: López, Francisco, et al.
Veröffentlicht: (2025)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
General Constrained Matrix Optimization
von: Garner, Casey, et al.
Veröffentlicht: (2024)
von: Garner, Casey, et al.
Veröffentlicht: (2024)
Fast and Accurate Interpolative Decompositions for General, Sparse, and Structured Tensors
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
On the Parenthesisations of Matrix Chains: All are Useful, Few Are Essential
von: López, Francisco, et al.
Veröffentlicht: (2023)
von: López, Francisco, et al.
Veröffentlicht: (2023)
Variability over 5 years of macrozoobenthos in the subtidal Wadden Sea near Sylt
von: Armonies, Werner
Veröffentlicht: (2017)
von: Armonies, Werner
Veröffentlicht: (2017)
Sub-Token Routing in LoRA for Adaptation and Query-Aware KV Compression
von: Jiang, Wei, et al.
Veröffentlicht: (2026)
von: Jiang, Wei, et al.
Veröffentlicht: (2026)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
Variability over 5 years of macrozoobenthos in the subtidal Wadden Sea near Sylt
von: Armonies, Werner
Veröffentlicht: (2017)
von: Armonies, Werner
Veröffentlicht: (2017)
Matrix-free stochastic calculation of operator norms without using adjoints
von: Bresch, Jonas, et al.
Veröffentlicht: (2024)
von: Bresch, Jonas, et al.
Veröffentlicht: (2024)
A simple polynomial-time approximation algorithm for the total variation distance between two product distributions
von: Feng, Weiming, et al.
Veröffentlicht: (2022)
von: Feng, Weiming, et al.
Veröffentlicht: (2022)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
Nutrients of pore water in sediments of the central equatorial Pacific
von: Martin, William R, et al.
Veröffentlicht: (1991)
von: Martin, William R, et al.
Veröffentlicht: (1991)
(Table 1) Climate data from the period 1963-2012 of 17 weather stations
von: De Jonge, Cindy, et al.
Veröffentlicht: (2014)
von: De Jonge, Cindy, et al.
Veröffentlicht: (2014)
Randomized Matrix Sketching for Neural Network Training and Gradient Monitoring
von: Antil, Harbir, et al.
Veröffentlicht: (2025)
von: Antil, Harbir, et al.
Veröffentlicht: (2025)
A Model Theoretic Perspective on Matrix Rings
von: Klep, Igor, et al.
Veröffentlicht: (2018)
von: Klep, Igor, et al.
Veröffentlicht: (2018)
Carbonate and Aluminium accumulation in the central equatorial Pacific Ocean
von: Murray, Richard W, et al.
Veröffentlicht: (1993)
von: Murray, Richard W, et al.
Veröffentlicht: (1993)
Part of the global DOC versus AOU (dissolved organic carbon/apparent oxygen utilization) data compilation, OACESWOCEP15S170W
von: Doval, María Dolores, et al.
Veröffentlicht: (2000)
von: Doval, María Dolores, et al.
Veröffentlicht: (2000)
SparseZipper: Enhancing Matrix Extensions to Accelerate SpGEMM on CPUs
von: Ta, Tuan, et al.
Veröffentlicht: (2025)
von: Ta, Tuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
von: Li, Zhonggen, et al.
Veröffentlicht: (2024) -
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026) -
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
von: Ma, Cong, et al.
Veröffentlicht: (2025) -
AES-SpMM: Balancing Accuracy and Speed by Adaptive Edge Sampling Strategy to Accelerate SpMM in GNNs
von: Song, Yingchen, et al.
Veröffentlicht: (2025) -
Hybrid SpMM SC26 AD
von: Wang, Haoyu
Veröffentlicht: (2026)