A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Jianhua, Liu, Bingjie, Ji, Weixing, Huang, Hua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
por: Gao, Jianhua, et al.
Publicado: (2024)
por: Gao, Jianhua, et al.
Publicado: (2024)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
por: Gao, Jianhua, et al.
Publicado: (2024)
por: Gao, Jianhua, et al.
Publicado: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
por: Böhnlein, Toni, et al.
Publicado: (2025)
por: Böhnlein, Toni, et al.
Publicado: (2025)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
por: Suzuki, Kengo, et al.
Publicado: (2026)
por: Suzuki, Kengo, et al.
Publicado: (2026)
Dynamic Memory Management on GPUs with SYCL
por: Standish, Russell K.
Publicado: (2025)
por: Standish, Russell K.
Publicado: (2025)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
por: Chillarón, Mónica, et al.
Publicado: (2024)
por: Chillarón, Mónica, et al.
Publicado: (2024)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
por: Ansari, Mufakir Qamar, et al.
Publicado: (2025)
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
por: Devarakonda, Aditya, et al.
Publicado: (2025)
por: Devarakonda, Aditya, et al.
Publicado: (2025)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
por: Cicirello, Vincent A.
Publicado: (2017)
por: Cicirello, Vincent A.
Publicado: (2017)
Stochastic well-structured transition systems
por: Aspnes, James
Publicado: (2025)
por: Aspnes, James
Publicado: (2025)
Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach
por: Xiaoteng, et al.
Publicado: (2024)
por: Xiaoteng, et al.
Publicado: (2024)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
por: Cheng, Long, et al.
Publicado: (2026)
por: Cheng, Long, et al.
Publicado: (2026)
Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion
por: Lappi, Oskar, et al.
Publicado: (2025)
por: Lappi, Oskar, et al.
Publicado: (2025)
Gradient Coding with Iterative Block Leverage Score Sampling
por: Charalambides, Neophytos, et al.
Publicado: (2023)
por: Charalambides, Neophytos, et al.
Publicado: (2023)
Distributed Hybrid Sketching for $\ell_2$-Embeddings
por: Charalambides, Neophytos, et al.
Publicado: (2024)
por: Charalambides, Neophytos, et al.
Publicado: (2024)
AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs
por: Li, Chendi, et al.
Publicado: (2022)
por: Li, Chendi, et al.
Publicado: (2022)
Declarative distributed algorithms as axiomatic theories in three-valued modal logic over semitopologies
por: Gabbay, Murdoch J.
Publicado: (2025)
por: Gabbay, Murdoch J.
Publicado: (2025)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
por: Kriemann, Ronald
Publicado: (2024)
por: Kriemann, Ronald
Publicado: (2024)
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
por: Mohammad, Noor Islam S.
Publicado: (2025)
por: Mohammad, Noor Islam S.
Publicado: (2025)
Algorithms for Parallel Shared-Memory Sparse Matrix-Vector Multiplication on Unstructured Matrices
por: Bergmans, Kobe, et al.
Publicado: (2025)
por: Bergmans, Kobe, et al.
Publicado: (2025)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
por: Zhao, Haisha, et al.
Publicado: (2025)
por: Zhao, Haisha, et al.
Publicado: (2025)
A simple protocol to automate the executing, scaling, and reconfiguration of Cloud-Native Apps
por: Ambroszkiewicz, Stanislaw, et al.
Publicado: (2023)
por: Ambroszkiewicz, Stanislaw, et al.
Publicado: (2023)
A Lock-Free, Fully GPU-Resident Architecture for the Verification of Goldbach's Conjecture
por: Llorente-Saguer, Isaac
Publicado: (2026)
por: Llorente-Saguer, Isaac
Publicado: (2026)
Distortion Resilience for Goal-Oriented Semantic Communication
por: Nguyen, Minh-Duong, et al.
Publicado: (2023)
por: Nguyen, Minh-Duong, et al.
Publicado: (2023)
Near-Optimal Bootstrapping of Hitting Sets for Algebraic Models
por: Kumar, Mrinal, et al.
Publicado: (2018)
por: Kumar, Mrinal, et al.
Publicado: (2018)
Stream parallel skeleton optimization
por: Aldinucci, Marco, et al.
Publicado: (2024)
por: Aldinucci, Marco, et al.
Publicado: (2024)
StreamFlow: cross-breeding cloud with HPC
por: Colonnelli, Iacopo, et al.
Publicado: (2020)
por: Colonnelli, Iacopo, et al.
Publicado: (2020)
A Study of Performance Portability in Plasma Physics Simulations
por: Ruzicka, Josef, et al.
Publicado: (2024)
por: Ruzicka, Josef, et al.
Publicado: (2024)
Anderson acceleration with approximate calculations: applications to scientific computing
por: Pasini, Massimiliano Lupo, et al.
Publicado: (2022)
por: Pasini, Massimiliano Lupo, et al.
Publicado: (2022)
Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems
por: Wang, Wenyi, et al.
Publicado: (2025)
por: Wang, Wenyi, et al.
Publicado: (2025)
Joint Training on AMD and NVIDIA GPUs
por: Hu, Jon, et al.
Publicado: (2026)
por: Hu, Jon, et al.
Publicado: (2026)
Enabling Practical Transparent Checkpointing for MPI: A Topological Sort Approach
por: Xu, Yao, et al.
Publicado: (2024)
por: Xu, Yao, et al.
Publicado: (2024)
Leveraging Multi-Instance GPUs through moldable task scheduling
por: Villarrubia, Jorge, et al.
Publicado: (2025)
por: Villarrubia, Jorge, et al.
Publicado: (2025)
Faster Linear Algebra Algorithms with Structured Random Matrices
por: Camaño, Chris, et al.
Publicado: (2025)
por: Camaño, Chris, et al.
Publicado: (2025)
Approximate Distributed Coded Computing: Polynomial Codes and Randomized Sketching
por: Charalambides, Neophytos, et al.
Publicado: (2026)
por: Charalambides, Neophytos, et al.
Publicado: (2026)
Minimum Cost Loop Nests for Contraction of a Sparse Tensor with a Tensor Network
por: Kanakagiri, Raghavendra, et al.
Publicado: (2023)
por: Kanakagiri, Raghavendra, et al.
Publicado: (2023)
Undercomplete Decomposition of Symmetric Tensors in Linear Time, and Smoothed Analysis of the Condition Number
por: Koiran, Pascal, et al.
Publicado: (2024)
por: Koiran, Pascal, et al.
Publicado: (2024)
NM-SpMM: Accelerating Matrix Multiplication Using N:M Sparsity with GPGPU
por: Ma, Cong, et al.
Publicado: (2025)
por: Ma, Cong, et al.
Publicado: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
por: Shakeri, Heman, et al.
Publicado: (2026)
por: Shakeri, Heman, et al.
Publicado: (2026)
Ejemplares similares
-
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
por: Gao, Jianhua, et al.
Publicado: (2024) -
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
por: Gao, Jianhua, et al.
Publicado: (2024) -
Efficient Parallel Scheduling for Sparse Triangular Solvers
por: Böhnlein, Toni, et al.
Publicado: (2025) -
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
por: Suzuki, Kengo, et al.
Publicado: (2026) -
Dynamic Memory Management on GPUs with SYCL
por: Standish, Russell K.
Publicado: (2025)