Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Shiju, Min, Younghoon, Yie, Hane, Kim, Hoshik, Ahn, Soohong, Sim, Joonseop, Lee, Chul-Ho, Kim, Jongryool |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
von: Yoon, Dongha, et al.
Veröffentlicht: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)
von: Li, Aiying, et al.
Veröffentlicht: (2026)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
HPU: High-Bandwidth Processing Unit for Scalable, Cost-effective LLM Inference via GPU Co-processing
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025)
SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks
von: Huang, Xin, et al.
Veröffentlicht: (2024)
von: Huang, Xin, et al.
Veröffentlicht: (2024)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
Improving Locality in Sparse and Dense Matrix Multiplications
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
von: Shah, Milan, et al.
Veröffentlicht: (2026)
von: Shah, Milan, et al.
Veröffentlicht: (2026)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
Accelerating high-order continuum kinetic plasma simulations using multiple GPUs
von: Ho, Andrew, et al.
Veröffentlicht: (2024)
von: Ho, Andrew, et al.
Veröffentlicht: (2024)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
von: Bellavita, Julian, et al.
Veröffentlicht: (2025)
von: Bellavita, Julian, et al.
Veröffentlicht: (2025)
DGEMM on Integer Matrix Multiplication Unit
von: Ootomo, Hiroyuki, et al.
Veröffentlicht: (2023)
von: Ootomo, Hiroyuki, et al.
Veröffentlicht: (2023)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
von: Remke, Stefan, et al.
Veröffentlicht: (2024)
von: Remke, Stefan, et al.
Veröffentlicht: (2024)
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
PARS3: Parallel Sparse Skew-Symmetric Matrix-Vector Multiplication with Reverse Cuthill-McKee Reordering
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Selection of Supervised Learning-based Sparse Matrix Reordering Algorithms
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
von: Hsieh, Chou-Ying, et al.
Veröffentlicht: (2024)
Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2024)
von: Gianinazzi, Lukas, et al.
Veröffentlicht: (2024)
Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
von: Adefemi, Temitayo
Veröffentlicht: (2024)
von: Adefemi, Temitayo
Veröffentlicht: (2024)
Demystifying ARM SME to Optimize General Matrix Multiplications
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
Fully-Automated Code Generation for Efficient Computation of Sparse Matrix Permanents on GPUs
von: Elbek, Deniz, et al.
Veröffentlicht: (2025)
von: Elbek, Deniz, et al.
Veröffentlicht: (2025)
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit
von: Uchino, Yuki, et al.
Veröffentlicht: (2024)
von: Uchino, Yuki, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
von: Yoon, Dongha, et al.
Veröffentlicht: (2025) -
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023) -
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference
von: Rhee, Myunghyun, et al.
Veröffentlicht: (2025) -
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026) -
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)