Selection of Supervised Learning-based Sparse Matrix Reordering Algorithms
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Tao, Jiang, Youfu, Cui, Yingbo, Fang, Jianbin, Zhang, Peng, Peng, Lin, Huang, Chun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
von: Asudeh, Omid, et al.
Veröffentlicht: (2025)
PARS3: Parallel Sparse Skew-Symmetric Matrix-Vector Multiplication with Reverse Cuthill-McKee Reordering
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
von: Yildirim, Selin, et al.
Veröffentlicht: (2024)
Demystifying ARM SME to Optimize General Matrix Multiplications
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
von: Deng, Chencheng, et al.
Veröffentlicht: (2025)
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024)
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
von: Islam, Abdullah Al Raqibul, et al.
Veröffentlicht: (2025)
von: Islam, Abdullah Al Raqibul, et al.
Veröffentlicht: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
von: Brock, Benjamin, et al.
Veröffentlicht: (2023)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
Sparsity-Aware Roofline Models for Sparse Matrix-Matrix Multiplication
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
von: Qian, Matthew, et al.
Veröffentlicht: (2026)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
von: Zhuang, Chen, et al.
Veröffentlicht: (2025)
ParamSpMM: Adaptive and Efficient Sparse Matrix-Matrix Multiplication on GPUs for GNNs
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
von: Zhang, Lixing, et al.
Veröffentlicht: (2026)
Accelerating Sparse Matrix-Matrix Multiplication on GPUs with Processing Near HBMs
von: Li, Shiju, et al.
Veröffentlicht: (2025)
von: Li, Shiju, et al.
Veröffentlicht: (2025)
MAGNUS: Generating Data Locality to Accelerate Sparse Matrix-Matrix Multiplication on CPUs
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
von: Wolfson-Pou, Jordi, et al.
Veröffentlicht: (2025)
Improving Locality in Sparse and Dense Matrix Multiplications
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
von: Dezfuli, Mohammad Mahdi Salehi, et al.
Veröffentlicht: (2024)
LOw-cOst yet High-Performant Sparse Matrix-Matrix Multiplication on Arm SME Architectures
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
von: Lei, Kelun, et al.
Veröffentlicht: (2025)
Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
von: Kong, Jie, et al.
Veröffentlicht: (2025)
von: Kong, Jie, et al.
Veröffentlicht: (2025)
HC-SpMM: Accelerating Sparse Matrix-Matrix Multiplication for Graphs with Hybrid GPU Cores
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
von: Li, Zhonggen, et al.
Veröffentlicht: (2024)
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3
von: Shah, Milan, et al.
Veröffentlicht: (2026)
von: Shah, Milan, et al.
Veröffentlicht: (2026)
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
von: Li, Aiying, et al.
Veröffentlicht: (2026)
von: Li, Aiying, et al.
Veröffentlicht: (2026)
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
von: Deng, Xiaoge, et al.
Veröffentlicht: (2021)
von: Deng, Xiaoge, et al.
Veröffentlicht: (2021)
Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
von: Adefemi, Temitayo
Veröffentlicht: (2024)
von: Adefemi, Temitayo
Veröffentlicht: (2024)
GPU-Accelerated Selected Basis Diagonalization with Thrust for SQD-based Algorithms
von: Doi, Jun, et al.
Veröffentlicht: (2026)
von: Doi, Jun, et al.
Veröffentlicht: (2026)
Efficient CPU-GPU Collaborative Inference for MoE-based LLMs on Memory-Limited Systems
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
von: Huang, En-Ming, et al.
Veröffentlicht: (2025)
Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations
von: Daas, Hussam Al, et al.
Veröffentlicht: (2024)
von: Daas, Hussam Al, et al.
Veröffentlicht: (2024)
A Survey of Distributed Graph Algorithms on Massive Graphs
von: Meng, Lingkai, et al.
Veröffentlicht: (2024)
von: Meng, Lingkai, et al.
Veröffentlicht: (2024)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
Matrix-PIC: Harnessing Matrix Outer-product for High-Performance Particle-in-Cell Simulations
von: Rao, Yizhuo, et al.
Veröffentlicht: (2026)
von: Rao, Yizhuo, et al.
Veröffentlicht: (2026)
A Structure-Aware Irregular Blocking Method for Sparse LU Factorization
von: Hu, Zhen, et al.
Veröffentlicht: (2025)
von: Hu, Zhen, et al.
Veröffentlicht: (2025)
Distributed Genetic Algorithm for Feature Selection
von: Potter, Michael, et al.
Veröffentlicht: (2024)
von: Potter, Michael, et al.
Veröffentlicht: (2024)
Lazy Qubit Reordering for Accelerating Parallel State-Vector-based Quantum Circuit Simulation
von: Teranishi, Yusuke, et al.
Veröffentlicht: (2024)
von: Teranishi, Yusuke, et al.
Veröffentlicht: (2024)
DAWN: Matrix Operation-Optimized Algorithm for Shortest Paths Problem on Unweighted Graphs
von: Feng, Yelai, et al.
Veröffentlicht: (2022)
von: Feng, Yelai, et al.
Veröffentlicht: (2022)
Fast Sparse Matrix Permutation for Mesh-Based Direct Solvers
von: Zarebavami, Behrooz, et al.
Veröffentlicht: (2026)
von: Zarebavami, Behrooz, et al.
Veröffentlicht: (2026)
RTop-K: Ultra-Fast Row-Wise Top-K Selection for Neural Network Acceleration on GPUs
von: Xie, Xi, et al.
Veröffentlicht: (2024)
von: Xie, Xi, et al.
Veröffentlicht: (2024)
Efficient Multi-Worker Selection based Distributed Swarm Learning via Analog Aggregation
von: Yao, Zhuoyu, et al.
Veröffentlicht: (2025)
von: Yao, Zhuoyu, et al.
Veröffentlicht: (2025)
Gensor: A Graph-based Construction Tensor Compilation Method for Deep Learning
von: Liu, Hangda, et al.
Veröffentlicht: (2025)
von: Liu, Hangda, et al.
Veröffentlicht: (2025)
Stencil Matrixization
von: Zhao, Wenxuan, et al.
Veröffentlicht: (2023)
von: Zhao, Wenxuan, et al.
Veröffentlicht: (2023)
Accelerating Sparse DNNs Based on Tiled GEMM
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
Efficient Gate Reordering for Distributed Quantum Compiling in Data Centers
von: Mengoni, Riccardo, et al.
Veröffentlicht: (2025)
von: Mengoni, Riccardo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
von: Asudeh, Omid, et al.
Veröffentlicht: (2025) -
PARS3: Parallel Sparse Skew-Symmetric Matrix-Vector Multiplication with Reverse Cuthill-McKee Reordering
von: Yildirim, Selin, et al.
Veröffentlicht: (2024) -
Demystifying ARM SME to Optimize General Matrix Multiplications
von: Deng, Chencheng, et al.
Veröffentlicht: (2025) -
Distributed-Memory Parallel Algorithms for Sparse Matrix and Sparse Tall-and-Skinny Matrix Multiplication
von: Ranawaka, Isuru, et al.
Veröffentlicht: (2024) -
Improving SpGEMM Performance Through Matrix Reordering and Cluster-wise Computation
von: Islam, Abdullah Al Raqibul, et al.
Veröffentlicht: (2025)