RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Brock, Benjamin, Buluç, Aydın, Yelick, Katherine |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BCL: A Cross-Platform Distributed Container Library
von: Brock, Benjamin, et al.
Veröffentlicht: (2018)
von: Brock, Benjamin, et al.
Veröffentlicht: (2018)
Distributed Matrix-Based Sampling for Graph Neural Network Training
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
Distributed-Memory Parallel Algorithms for Fixed-Radius Near Neighbor Graph Construction
von: Raulet, Gabriel, et al.
Veröffentlicht: (2025)
von: Raulet, Gabriel, et al.
Veröffentlicht: (2025)
The Ubiquitous Sparse Matrix-Matrix Products
von: Buluç, Aydın
Veröffentlicht: (2025)
von: Buluç, Aydın
Veröffentlicht: (2025)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
von: Li, Shiju, et al.
Veröffentlicht: (2025)
von: Li, Shiju, et al.
Veröffentlicht: (2025)
A sparsity-aware distributed-memory algorithm for sparse-sparse matrix multiplication
von: Hong, Yuxi, et al.
Veröffentlicht: (2024)
von: Hong, Yuxi, et al.
Veröffentlicht: (2024)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
von: Islam, Abdullah Al Raqibul, et al.
Veröffentlicht: (2025)
von: Islam, Abdullah Al Raqibul, et al.
Veröffentlicht: (2025)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
Scaling Graph Neural Networks for Particle Track Reconstruction
von: Tripathy, Alok, et al.
Veröffentlicht: (2025)
von: Tripathy, Alok, et al.
Veröffentlicht: (2025)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)
von: Li, Aiying, et al.
Veröffentlicht: (2026)
Slicing Is All You Need: Towards A Universal One-Sided Algorithm for Distributed Matrix Multiplication
von: Brock, Benjamin, et al.
Veröffentlicht: (2025)
von: Brock, Benjamin, et al.
Veröffentlicht: (2025)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
Fast Algorithms for Scheduling Many-body Correlation Functions on Accelerators
von: Selvitopi, Oguz, et al.
Veröffentlicht: (2025)
von: Selvitopi, Oguz, et al.
Veröffentlicht: (2025)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
Improving Locality in Sparse and Dense Matrix Multiplications
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
ALock: Asymmetric Lock Primitive for RDMA Systems
von: Baran, Amanda, et al.
Veröffentlicht: (2024)
von: Baran, Amanda, et al.
Veröffentlicht: (2024)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
Towards Efficient and Scalable Distributed Vector Search with RDMA
von: Zhi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhi, Xiangyu, et al.
Veröffentlicht: (2025)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
von: Shah, Milan, et al.
Veröffentlicht: (2026)
von: Shah, Milan, et al.
Veröffentlicht: (2026)
Parallelizing the Approximate Minimum Degree Ordering Algorithm: Strategies and Evaluation
von: Chang, Yen-Hsiang, et al.
Veröffentlicht: (2025)
von: Chang, Yen-Hsiang, et al.
Veröffentlicht: (2025)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
fabric-lib: RDMA Point-to-Point Communication for LLM Systems
von: Licker, Nandor, et al.
Veröffentlicht: (2025)
von: Licker, Nandor, et al.
Veröffentlicht: (2025)
The Semantic Arrow of Time, Part III: RDMA and the Completion Fallacy
von: Borrill, Paul
Veröffentlicht: (2026)
von: Borrill, Paul
Veröffentlicht: (2026)
Selection of Supervised Learning-based Sparse Matrix Reordering Algorithms
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
von: Adefemi, Temitayo
Veröffentlicht: (2024)
von: Adefemi, Temitayo
Veröffentlicht: (2024)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
von: Dwaraknath, Rajat Vadiraj, et al.
Veröffentlicht: (2026)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
von: Bellavita, Julian, et al.
Veröffentlicht: (2025)
von: Bellavita, Julian, et al.
Veröffentlicht: (2025)
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU
von: Li, Yifan, et al.
Veröffentlicht: (2026)
von: Li, Yifan, et al.
Veröffentlicht: (2026)
CB-SpMV:A Data Aggregating and Balance Algorithm for Cache-Friendly Block-Based SpMV on GPUs
von: Cong, Xing, et al.
Veröffentlicht: (2026)
von: Cong, Xing, et al.
Veröffentlicht: (2026)
PARS3: Parallel Sparse Skew-Symmetric Matrix-Vector Multiplication with Reverse Cuthill-McKee Reordering
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
Efficiently Parallelizable Strassen-Based Multiplication of a Matrix by its Transpose
von: Arrigoni, Viviana, et al.
Veröffentlicht: (2021)
von: Arrigoni, Viviana, et al.
Veröffentlicht: (2021)
CPMA: An Efficient Batch-Parallel Compressed Set Without Pointers
von: Wheatman, Brian, et al.
Veröffentlicht: (2023)
von: Wheatman, Brian, et al.
Veröffentlicht: (2023)
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
von: Chen, June, et al.
Veröffentlicht: (2026)
von: Chen, June, et al.
Veröffentlicht: (2026)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
von: Zhang, Zhekai, et al.
Veröffentlicht: (2020)
Ähnliche Einträge
-
BCL: A Cross-Platform Distributed Container Library
von: Brock, Benjamin, et al.
Veröffentlicht: (2018) -
Distributed Matrix-Based Sampling for Graph Neural Network Training
von: Tripathy, Alok, et al.
Veröffentlicht: (2023) -
Distributed-Memory Parallel Algorithms for Fixed-Radius Near Neighbor Graph Construction
von: Raulet, Gabriel, et al.
Veröffentlicht: (2025) -
The Ubiquitous Sparse Matrix-Matrix Products
von: Buluç, Aydın
Veröffentlicht: (2025) -
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
von: Li, Shiju, et al.
Veröffentlicht: (2025)