Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Adefemi, Temitayo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do MPI Derived Datatypes Actually Help? A Single-Node Cross-Implementation Study on Shared-Memory Communication
von: Adefemi, Temitayo
Veröffentlicht: (2025)
von: Adefemi, Temitayo
Veröffentlicht: (2025)
The Entropy of Parallel Systems
von: Adefemi, Temitayo
Veröffentlicht: (2025)
von: Adefemi, Temitayo
Veröffentlicht: (2025)
What Every Computer Scientist Needs To Know About Parallelization
von: Adefemi, Temitayo
Veröffentlicht: (2025)
von: Adefemi, Temitayo
Veröffentlicht: (2025)
Otus Supercomputer
von: Ehtesabi, Sadaf, et al.
Veröffentlicht: (2025)
von: Ehtesabi, Sadaf, et al.
Veröffentlicht: (2025)
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
von: Shubham, et al.
Veröffentlicht: (2024)
von: Shubham, et al.
Veröffentlicht: (2024)
TX-Digital Twin: Visualizing Supercomputer GPU Performance Data Stream
von: Baskakova, Elena, et al.
Veröffentlicht: (2026)
von: Baskakova, Elena, et al.
Veröffentlicht: (2026)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit
von: Uchino, Yuki, et al.
Veröffentlicht: (2024)
von: Uchino, Yuki, et al.
Veröffentlicht: (2024)
Enabling Message Passing Interface Containers on the LUMI Supercomputer
von: Lazzaro, Alfio
Veröffentlicht: (2024)
von: Lazzaro, Alfio
Veröffentlicht: (2024)
High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
von: Jangda, Abhinav, et al.
Veröffentlicht: (2024)
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
von: Takahashi, Keichi, et al.
Veröffentlicht: (2023)
MalleTrain: Deep Neural Network Training on Unfillable Supercomputer Nodes
von: Ma, Xiaolong, et al.
Veröffentlicht: (2024)
von: Ma, Xiaolong, et al.
Veröffentlicht: (2024)
Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
von: Kinkead, Shannon, et al.
Veröffentlicht: (2026)
von: Kinkead, Shannon, et al.
Veröffentlicht: (2026)
Supercomputer 3D Digital Twin for User Focused Real-Time Monitoring
von: Bergeron, William, et al.
Veröffentlicht: (2024)
von: Bergeron, William, et al.
Veröffentlicht: (2024)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
DGEMM on Integer Matrix Multiplication Unit
von: Ootomo, Hiroyuki, et al.
Veröffentlicht: (2023)
von: Ootomo, Hiroyuki, et al.
Veröffentlicht: (2023)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
von: Li, Shiju, et al.
Veröffentlicht: (2025)
von: Li, Shiju, et al.
Veröffentlicht: (2025)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
Supercomputing for High-speed Avoidance and Reactive Planning in Robots
von: Lachmansingh, Kieran S., et al.
Veröffentlicht: (2025)
von: Lachmansingh, Kieran S., et al.
Veröffentlicht: (2025)
Improving Locality in Sparse and Dense Matrix Multiplications
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension
von: Remke, Stefan, et al.
Veröffentlicht: (2024)
von: Remke, Stefan, et al.
Veröffentlicht: (2024)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
Demystifying ARM SME to Optimize General Matrix Multiplications
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
PWDFT-SW: Extending the Limit of Plane-Wave DFT Calculations to 16K Atoms on the New Sunway Supercomputer
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
von: Jiang, Qingcai, et al.
Veröffentlicht: (2024)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)
von: Li, Aiying, et al.
Veröffentlicht: (2026)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
von: Shah, Milan, et al.
Veröffentlicht: (2026)
von: Shah, Milan, et al.
Veröffentlicht: (2026)
Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
von: Uchino, Yuki, et al.
Veröffentlicht: (2025)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
von: Cornelius, Melanie, et al.
Veröffentlicht: (2025)
Tolerance to Asynchrony in Algorithms for Multiplication and Modulo
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
von: Gupta, Arya Tanmay, et al.
Veröffentlicht: (2023)
Matrix-PIC: Harnessing Matrix Outer-product for High-Performance Particle-in-Cell Simulations
von: Rao, Yizhuo, et al.
Veröffentlicht: (2026)
von: Rao, Yizhuo, et al.
Veröffentlicht: (2026)
Efficiently Parallelizable Strassen-Based Multiplication of a Matrix by its Transpose
von: Arrigoni, Viviana, et al.
Veröffentlicht: (2021)
von: Arrigoni, Viviana, et al.
Veröffentlicht: (2021)
Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations
von: Daas, Hussam Al, et al.
Veröffentlicht: (2024)
von: Daas, Hussam Al, et al.
Veröffentlicht: (2024)
Selection of Supervised Learning-based Sparse Matrix Reordering Algorithms
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
Comparative Analysis of Distributed Caching Algorithms: Performance Metrics and Implementation Considerations
von: Mayer, Helen, et al.
Veröffentlicht: (2025)
von: Mayer, Helen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do MPI Derived Datatypes Actually Help? A Single-Node Cross-Implementation Study on Shared-Memory Communication
von: Adefemi, Temitayo
Veröffentlicht: (2025) -
The Entropy of Parallel Systems
von: Adefemi, Temitayo
Veröffentlicht: (2025) -
What Every Computer Scientist Needs To Know About Parallelization
von: Adefemi, Temitayo
Veröffentlicht: (2025) -
Otus Supercomputer
von: Ehtesabi, Sadaf, et al.
Veröffentlicht: (2025) -
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
von: Shubham, et al.
Veröffentlicht: (2024)