SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhuang, Chen, Zhang, Lingqi, Brock, Benjamin, Wu, Du, Chen, Peng, Endo, Toshio, Matsuoka, Satoshi, Wahib, Mohamed |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
von: Lacey, Dane C., et al.
Veröffentlicht: (2024)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
A Multi-Port Concurrent Communication Model for handling Compute Intensive Tasks on Distributed Satellite System Constellations
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
von: Veeravalli, Bharadwaj
Veröffentlicht: (2026)
Staging Blocked Evaluation over Structured Sparse Matrices
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
von: Das, Pratyush, et al.
Veröffentlicht: (2024)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
von: Karfakis, George, et al.
Veröffentlicht: (2025)
von: Karfakis, George, et al.
Veröffentlicht: (2025)
Modeling the Effect of Data Redundancy on Speedup in MLFMA Near-Field Computation
von: Sadeghi, Morteza
Veröffentlicht: (2025)
von: Sadeghi, Morteza
Veröffentlicht: (2025)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
von: Chang, Dali, et al.
Veröffentlicht: (2026)
von: Chang, Dali, et al.
Veröffentlicht: (2026)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
von: Xu, Jingwei, et al.
Veröffentlicht: (2025)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
von: Suffa, Philipp, et al.
Veröffentlicht: (2024)
Optimal Parallel Scheduling under Concave Speedup Functions
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
von: Li, Chengzhang, et al.
Veröffentlicht: (2025)
Optimal Configuration of API Resources in Cloud Native Computing
von: Truyen, Eddy, et al.
Veröffentlicht: (2025)
von: Truyen, Eddy, et al.
Veröffentlicht: (2025)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
von: Sheng, Peiyao, et al.
Veröffentlicht: (2024)
An Online Probabilistic Distributed Tracing System
von: Toslali, M., et al.
Veröffentlicht: (2024)
von: Toslali, M., et al.
Veröffentlicht: (2024)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
Operational Strategies for Non-Disruptive Scheduling Transitions in Production HPC Systems
von: MacLachlan, Glen, et al.
Veröffentlicht: (2026)
von: MacLachlan, Glen, et al.
Veröffentlicht: (2026)
Distributed Matrix-Based Sampling for Graph Neural Network Training
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
von: Tripathy, Alok, et al.
Veröffentlicht: (2023)
Kubernetes in Action: Exploring the Performance of Kubernetes Distributions in the Cloud
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024)
Ridgeline: A 2D Roofline Model for Distributed Systems
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
von: Checconi, Fabio, et al.
Veröffentlicht: (2022)
Communication-Aware Diffusion Load Balancing for Persistently Interacting Objects
von: Taylor, Maya, et al.
Veröffentlicht: (2026)
von: Taylor, Maya, et al.
Veröffentlicht: (2026)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
von: Owen, Herbert, et al.
Veröffentlicht: (2024)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
von: Vatsavai, Sairam Sri, et al.
Veröffentlicht: (2025)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
von: McDonald, Jesse, et al.
Veröffentlicht: (2024)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
von: Shan, Baodi, et al.
Veröffentlicht: (2024)
Near-Optimal Wafer-Scale Reduce
von: Luczynski, Piotr, et al.
Veröffentlicht: (2024)
von: Luczynski, Piotr, et al.
Veröffentlicht: (2024)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
How to Rent GPUs on a Budget
von: Li, Zhouzi, et al.
Veröffentlicht: (2024)
von: Li, Zhouzi, et al.
Veröffentlicht: (2024)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
von: Chen, David, et al.
Veröffentlicht: (2026)
von: Chen, David, et al.
Veröffentlicht: (2026)
Advanced Scheduling Strategies for Distributed Quantum Computing Jobs
von: Ni, Gongyu, et al.
Veröffentlicht: (2026)
von: Ni, Gongyu, et al.
Veröffentlicht: (2026)
BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
von: Hu, Xiannan, et al.
Veröffentlicht: (2025)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
von: Lin, Wei-Chen, et al.
Veröffentlicht: (2024)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
von: Ather, Hammad, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024) -
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
von: Zhang, Lingqi, et al.
Veröffentlicht: (2025) -
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025) -
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
von: Katagiri, Takahiro, et al.
Veröffentlicht: (2024) -
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)