HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhonggen, Ke, Xiangyu, Zhu, Yifan, Gao, Yunjun, Tu, Yaofeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)
von: Li, Aiying, et al.
Veröffentlicht: (2026)
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
Accelerating Biclique Counting on GPU
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
AES-SpMM: Balancing Accuracy and Speed by Adaptive Edge Sampling Strategy to Accelerate SpMM in GNNs
von: Song, Yingchen, et al.
Veröffentlicht: (2025)
von: Song, Yingchen, et al.
Veröffentlicht: (2025)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
High Performance Unstructured SpMM Computation Using Tensor Cores
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
von: Okanovic, Patrik, et al.
Veröffentlicht: (2024)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
von: Ma, Cong, et al.
Veröffentlicht: (2025)
von: Ma, Cong, et al.
Veröffentlicht: (2025)
Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor Search
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
von: Li, Zhonggen, et al.
Veröffentlicht: (2025)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
von: Li, Bohan, et al.
Veröffentlicht: (2026)
von: Li, Bohan, et al.
Veröffentlicht: (2026)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
von: Li, Shiju, et al.
Veröffentlicht: (2025)
von: Li, Shiju, et al.
Veröffentlicht: (2025)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
Improving Locality in Sparse and Dense Matrix Multiplications
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
von: Shah, Milan, et al.
Veröffentlicht: (2026)
von: Shah, Milan, et al.
Veröffentlicht: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
von: McFarland, Thomas, et al.
Veröffentlicht: (2025)
GPU Accelerated Sparse Cholesky Factorization
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
von: Karsavuran, M. Ozan, et al.
Veröffentlicht: (2024)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
von: Islam, Abdullah Al Raqibul, et al.
Veröffentlicht: (2025)
von: Islam, Abdullah Al Raqibul, et al.
Veröffentlicht: (2025)
Matrix representation and GPU-optimized parallel B-spline computing
von: Wu, Jiayu, et al.
Veröffentlicht: (2025)
von: Wu, Jiayu, et al.
Veröffentlicht: (2025)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
DGEMM on Integer Matrix Multiplication Unit
von: Ootomo, Hiroyuki, et al.
Veröffentlicht: (2023)
von: Ootomo, Hiroyuki, et al.
Veröffentlicht: (2023)
GPU-Accelerated Vecchia Approximations of Gaussian Processes for Geospatial Data using Batched Matrix Computations
von: Pan, Qilong, et al.
Veröffentlicht: (2024)
von: Pan, Qilong, et al.
Veröffentlicht: (2024)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
von: Remke, Stefan, et al.
Veröffentlicht: (2024)
von: Remke, Stefan, et al.
Veröffentlicht: (2024)
PICO: Accelerating All k-Core Paradigms on GPU
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
PARS3: Parallel Sparse Skew-Symmetric Matrix-Vector Multiplication with Reverse Cuthill-McKee Reordering
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
MERBIT: A GPU-Based SpMV Method for Iterative Workloads
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
von: Zhang, Qi, et al.
Veröffentlicht: (2026)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
von: Schieffer, Gabin, et al.
Veröffentlicht: (2024)
Selection of Supervised Learning-based Sparse Matrix Reordering Algorithms
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
von: Zhao, Haisha, et al.
Veröffentlicht: (2025) -
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026) -
Efficient Graph Embedding at Scale: Optimizing CPU-GPU-SSD Integration
von: Li, Zhonggen, et al.
Veröffentlicht: (2025) -
Accelerating Biclique Counting on GPU
von: Qiu, Linshan, et al.
Veröffentlicht: (2024) -
AES-SpMM: Balancing Accuracy and Speed by Adaptive Edge Sampling Strategy to Accelerate SpMM in GNNs
von: Song, Yingchen, et al.
Veröffentlicht: (2025)