SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
Fuente:
arXiv
Saved in:
| Main Authors: | Abubaker, Nabil, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
by: Gianinazzi, Lukas, et al.
Published: (2024)
by: Gianinazzi, Lukas, et al.
Published: (2024)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
by: Li, Aiying, et al.
Published: (2026)
by: Li, Aiying, et al.
Published: (2026)
High Performance Unstructured SpMM Computation Using Tensor Cores
by: Okanovic, Patrik, et al.
Published: (2024)
by: Okanovic, Patrik, et al.
Published: (2024)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026)
by: Shah, Milan, et al.
Published: (2026)
Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage
by: Lin, Junqing, et al.
Published: (2025)
by: Lin, Junqing, et al.
Published: (2025)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
by: Zhang, Lixing, et al.
Published: (2026)
by: Zhang, Lixing, et al.
Published: (2026)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
by: Li, Zhonggen, et al.
Published: (2024)
by: Li, Zhonggen, et al.
Published: (2024)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
by: Zhang, Zhekai, et al.
Published: (2020)
by: Zhang, Zhekai, et al.
Published: (2020)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
by: McFarland, Thomas, et al.
Published: (2025)
by: McFarland, Thomas, et al.
Published: (2025)
Near-Optimal Sparse Allreduce for Distributed Deep Learning
by: Li, Shigang, et al.
Published: (2022)
by: Li, Shigang, et al.
Published: (2022)
FlashSketch: Sketch-Kernel Co-Design for Fast Sparse Sketching on GPUs
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
by: Dwaraknath, Rajat Vadiraj, et al.
Published: (2026)
Popcorn: Accelerating Kernel K-means on GPUs through Sparse Linear Algebra
by: Bellavita, Julian, et al.
Published: (2025)
by: Bellavita, Julian, et al.
Published: (2025)
Cppless: Single-Source and High-Performance Serverless Programming in C++
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
by: Chen, Tiancheng, et al.
Published: (2025)
by: Chen, Tiancheng, et al.
Published: (2025)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
by: Lepori, Andrea, et al.
Published: (2025)
by: Lepori, Andrea, et al.
Published: (2025)
BLaST: High Performance Inference and Pretraining using BLock Sparse Transformers
by: Okanovic, Patrik, et al.
Published: (2025)
by: Okanovic, Patrik, et al.
Published: (2025)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
by: Noh, Si Ung, et al.
Published: (2024)
by: Noh, Si Ung, et al.
Published: (2024)
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example
by: Copik, Marcin, et al.
Published: (2022)
by: Copik, Marcin, et al.
Published: (2022)
Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
by: Deng, Xiaoge, et al.
Published: (2021)
by: Deng, Xiaoge, et al.
Published: (2021)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
by: Fusco, Luigi, et al.
Published: (2024)
by: Fusco, Luigi, et al.
Published: (2024)
SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array
by: Ni, Xiaobing, et al.
Published: (2024)
by: Ni, Xiaobing, et al.
Published: (2024)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
by: Ranawaka, Isuru, et al.
Published: (2024)
by: Ranawaka, Isuru, et al.
Published: (2024)
SparseServe: Unlocking Parallelism for Dynamic Sparse Attention in Long-Context LLM Serving
by: Zhou, Qihui, et al.
Published: (2025)
by: Zhou, Qihui, et al.
Published: (2025)
XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing
by: Hoefler, Torsten, et al.
Published: (2024)
by: Hoefler, Torsten, et al.
Published: (2024)
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation
by: Chen, Fahao, et al.
Published: (2024)
by: Chen, Fahao, et al.
Published: (2024)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
by: Qing, Yuhao, et al.
Published: (2025)
by: Qing, Yuhao, et al.
Published: (2025)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
by: Yan, Ran, et al.
Published: (2025)
by: Yan, Ran, et al.
Published: (2025)
Inductive Loop Analysis for Practical HPC Application Optimization
by: Schaad, Philipp, et al.
Published: (2025)
by: Schaad, Philipp, et al.
Published: (2025)
GPU Accelerated Sparse Cholesky Factorization
by: Karsavuran, M. Ozan, et al.
Published: (2024)
by: Karsavuran, M. Ozan, et al.
Published: (2024)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
by: Khalilov, Mikhail, et al.
Published: (2024)
by: Khalilov, Mikhail, et al.
Published: (2024)
A More Scalable Sparse Dynamic Data Exchange
by: Geyko, Andrew, et al.
Published: (2023)
by: Geyko, Andrew, et al.
Published: (2023)
CB-SpMV:A Data Aggregating and Balance Algorithm for Cache-Friendly Block-Based SpMV on GPUs
by: Cong, Xing, et al.
Published: (2026)
by: Cong, Xing, et al.
Published: (2026)
Federated LoRA with Sparse Communication
by: Kuo, Kevin, et al.
Published: (2024)
by: Kuo, Kevin, et al.
Published: (2024)
Load Balancing Using Sparse Communication
by: Mendelson, Gal, et al.
Published: (2022)
by: Mendelson, Gal, et al.
Published: (2022)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
by: Asudeh, Omid, et al.
Published: (2025)
by: Asudeh, Omid, et al.
Published: (2025)
Accelerating Sparse DNNs Based on Tiled GEMM
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
Improving Locality in Sparse and Dense Matrix Multiplications
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
by: Dezfuli, Mohammad Mahdi Salehi, et al.
Published: (2024)
Similar Items
-
Arrow Matrix Decomposition: A Novel Approach for Communication-Efficient Sparse Matrix Multiplication
by: Gianinazzi, Lukas, et al.
Published: (2024) -
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
by: Li, Aiying, et al.
Published: (2026) -
High Performance Unstructured SpMM Computation Using Tensor Cores
by: Okanovic, Patrik, et al.
Published: (2024) -
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
by: Shah, Milan, et al.
Published: (2026) -
Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage
by: Lin, Junqing, et al.
Published: (2025)