A sparsity-aware distributed-memory algorithm for sparse-sparse matrix multiplication
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Yuxi, Buluc, Aydin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast multiplication of random dense matrices with fixed sparse matrices
by: Liang, Tianyu, et al.
Published: (2023)
by: Liang, Tianyu, et al.
Published: (2023)
BCL: A Cross-Platform Distributed Container Library
by: Brock, Benjamin, et al.
Published: (2018)
by: Brock, Benjamin, et al.
Published: (2018)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025)
The Ubiquitous Sparse Matrix-Matrix Products
by: Buluç, Aydın
Published: (2025)
by: Buluç, Aydın
Published: (2025)
Distributed Matrix-Based Sampling for Graph Neural Network Training
by: Tripathy, Alok, et al.
Published: (2023)
by: Tripathy, Alok, et al.
Published: (2023)
Fast Algorithms for Scheduling Many-body Correlation Functions on Accelerators
by: Selvitopi, Oguz, et al.
Published: (2025)
by: Selvitopi, Oguz, et al.
Published: (2025)
On the energy efficiency of sparse matrix computations on multi-GPU clusters
by: Bernaschi, Massimo, et al.
Published: (2025)
by: Bernaschi, Massimo, et al.
Published: (2025)
CPMA: An Efficient Batch-Parallel Compressed Set Without Pointers
by: Wheatman, Brian, et al.
Published: (2023)
by: Wheatman, Brian, et al.
Published: (2023)
Distributed-Memory Parallel Algorithms for Fixed-Radius Near Neighbor Graph Construction
by: Raulet, Gabriel, et al.
Published: (2025)
by: Raulet, Gabriel, et al.
Published: (2025)
Parallelizing the Approximate Minimum Degree Ordering Algorithm: Strategies and Evaluation
by: Chang, Yen-Hsiang, et al.
Published: (2025)
by: Chang, Yen-Hsiang, et al.
Published: (2025)
Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors
by: Chen, Siyuan, et al.
Published: (2024)
by: Chen, Siyuan, et al.
Published: (2024)
Evaluation of computational and energy performance in matrix multiplication algorithms on CPU and GPU using MKL, cuBLAS and SYCL
by: Torres, L. A., et al.
Published: (2024)
by: Torres, L. A., et al.
Published: (2024)
Scaling atomic ordering in shared memory
by: Martignetti, Lorenzo, et al.
Published: (2025)
by: Martignetti, Lorenzo, et al.
Published: (2025)
Scalability of 3D-DFT by block tensor-matrix multiplication on the JUWELS Cluster
by: Malapally, Nitin, et al.
Published: (2023)
by: Malapally, Nitin, et al.
Published: (2023)
Affinity-aware Serverless Function Scheduling
by: De Palma, Giuseppe, et al.
Published: (2024)
by: De Palma, Giuseppe, et al.
Published: (2024)
emucxl: an emulation framework for CXL-based disaggregated memory applications
by: Gond, Raja, et al.
Published: (2024)
by: Gond, Raja, et al.
Published: (2024)
A framework to reason about consistency and atomicity guarantees in a sparsely-connected, partially-replicated peer-to-peer system
by: Nair, Sreeja S., et al.
Published: (2026)
by: Nair, Sreeja S., et al.
Published: (2026)
Energy-aware operation of HPC systems in Germany
by: Suarez, Estela, et al.
Published: (2024)
by: Suarez, Estela, et al.
Published: (2024)
A shared compilation stack for distributed-memory parallelism in stencil DSLs
by: Bisbas, George, et al.
Published: (2024)
by: Bisbas, George, et al.
Published: (2024)
Scaling Graph Neural Networks for Particle Track Reconstruction
by: Tripathy, Alok, et al.
Published: (2025)
by: Tripathy, Alok, et al.
Published: (2025)
LRScheduler: A Layer-aware and Resource-adaptive Container Scheduler in Edge Computing
by: Tang, Zhiqing, et al.
Published: (2025)
by: Tang, Zhiqing, et al.
Published: (2025)
Energy-aware Distributed Microservice Request Placement at the Edge
by: Toczé, Klervie, et al.
Published: (2024)
by: Toczé, Klervie, et al.
Published: (2024)
Dependency-aware Resource Allocation for Serverless Functions at the Edge
by: Baresi, Luciano, et al.
Published: (2023)
by: Baresi, Luciano, et al.
Published: (2023)
DUMBO: Making durable read-only transactions fly on hardware transactional memory
by: Barreto, João, et al.
Published: (2024)
by: Barreto, João, et al.
Published: (2024)
Memory-aware Adaptive Scheduling of Scientific Workflows on Heterogeneous Architectures
by: Kulagina, Svetlana, et al.
Published: (2025)
by: Kulagina, Svetlana, et al.
Published: (2025)
Truncated multiplication and batch software SIMD AVX512 implementation for faster Montgomery multiplications and modular exponentiation
by: Didier, Laurent-Stéphane, et al.
Published: (2024)
by: Didier, Laurent-Stéphane, et al.
Published: (2024)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
by: Duan, Jiaang, et al.
Published: (2025)
by: Duan, Jiaang, et al.
Published: (2025)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
by: Wen, Linfeng, et al.
Published: (2024)
by: Wen, Linfeng, et al.
Published: (2024)
An Interference-aware Approach for Co-located Container Orchestration with Novel Metric
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
by: Luo, Shuqing, et al.
Published: (2024)
by: Luo, Shuqing, et al.
Published: (2024)
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
by: Hanafy, Walid A., et al.
Published: (2025)
by: Hanafy, Walid A., et al.
Published: (2025)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
by: Xuan, Mo, et al.
Published: (2025)
by: Xuan, Mo, et al.
Published: (2025)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
by: Hua, Qin, et al.
Published: (2024)
by: Hua, Qin, et al.
Published: (2024)
A Two-Level Thermal Cycling-aware Task Mapping Technique for Reliability Management in Manycore Systems
by: Khani, Fatemeh Hossein, et al.
Published: (2024)
by: Khani, Fatemeh Hossein, et al.
Published: (2024)
A Framework for Carbon-aware Real-Time Workload Management in Clouds using Renewables-driven Cores
by: Hewage, Tharindu B., et al.
Published: (2024)
by: Hewage, Tharindu B., et al.
Published: (2024)
Tangram: High-resolution Video Analytics on Serverless Platform with SLO-aware Batching
by: Peng, Haosong, et al.
Published: (2024)
by: Peng, Haosong, et al.
Published: (2024)
vPALs: Towards Verified Performance-aware Learning System For Resource Management
by: He, Guoliang, et al.
Published: (2024)
by: He, Guoliang, et al.
Published: (2024)
Energy-aware Incremental OTA Update for Flash-based Batteryless IoT Devices
by: Wei, Wei, et al.
Published: (2024)
by: Wei, Wei, et al.
Published: (2024)
Workflow decomposition algorithm for scheduling with quantum annealer-based hybrid solver
by: Kroczek, Marcin, et al.
Published: (2025)
by: Kroczek, Marcin, et al.
Published: (2025)
Similar Items
-
Fast multiplication of random dense matrices with fixed sparse matrices
by: Liang, Tianyu, et al.
Published: (2023) -
BCL: A Cross-Platform Distributed Container Library
by: Brock, Benjamin, et al.
Published: (2018) -
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023) -
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
by: Islam, Abdullah Al Raqibul, et al.
Published: (2025) -
The Ubiquitous Sparse Matrix-Matrix Products
by: Buluç, Aydın
Published: (2025)