Algorithms for Parallel Shared-Memory Sparse Matrix-Vector Multiplication on Unstructured Matrices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bergmans, Kobe, Meerbergen, Karl, Vandebril, Raf |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
Real-time chaotic video encryption based on multithreaded parallel confusion and diffusion
von: Jiang, Dong, et al.
Veröffentlicht: (2023)
von: Jiang, Dong, et al.
Veröffentlicht: (2023)
GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN
von: Yin, Jia, et al.
Veröffentlicht: (2025)
von: Yin, Jia, et al.
Veröffentlicht: (2025)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
von: Kriemann, Ronald
Veröffentlicht: (2024)
von: Kriemann, Ronald
Veröffentlicht: (2024)
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
Dynamic Memory Management on GPUs with SYCL
von: Standish, Russell K.
Veröffentlicht: (2025)
von: Standish, Russell K.
Veröffentlicht: (2025)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
Population Protocols Revisited: Parity and Beyond
von: Gąsieniec, Leszek, et al.
Veröffentlicht: (2025)
von: Gąsieniec, Leszek, et al.
Veröffentlicht: (2025)
OPTIMUM-DERAM: Highly Consistent, Scalable, and Secure Multi-Object Memory using RLNC
von: Nicolaou, Nicolas, et al.
Veröffentlicht: (2026)
von: Nicolaou, Nicolas, et al.
Veröffentlicht: (2026)
A Morton-Type Space-Filling Curve for Pyramid Subdivision and Hybrid Adaptive Mesh Refinement
von: Knapp, David, et al.
Veröffentlicht: (2026)
von: Knapp, David, et al.
Veröffentlicht: (2026)
Faster Vertex Cover Algorithms on GPUs with Component-Aware Parallel Branching
von: Amro, Hussein, et al.
Veröffentlicht: (2025)
von: Amro, Hussein, et al.
Veröffentlicht: (2025)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
von: Chillarón, Mónica, et al.
Veröffentlicht: (2024)
von: Chillarón, Mónica, et al.
Veröffentlicht: (2024)
Two parallel dynamic lexicographic algorithms for factorization sets in numerical semigroups
von: Barron, Thomas
Veröffentlicht: (2024)
von: Barron, Thomas
Veröffentlicht: (2024)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
Context Adaptive Cooperation
von: Albouy, Timothé, et al.
Veröffentlicht: (2023)
von: Albouy, Timothé, et al.
Veröffentlicht: (2023)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
von: Cicirello, Vincent A.
Veröffentlicht: (2017)
von: Cicirello, Vincent A.
Veröffentlicht: (2017)
Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion
von: Lappi, Oskar, et al.
Veröffentlicht: (2025)
von: Lappi, Oskar, et al.
Veröffentlicht: (2025)
Machine Learning-Driven Predictive Resource Management in Complex Science Workflows
von: Chowdhury, Tasnuva, et al.
Veröffentlicht: (2025)
von: Chowdhury, Tasnuva, et al.
Veröffentlicht: (2025)
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
von: Ferreira, Inês, et al.
Veröffentlicht: (2024)
von: Ferreira, Inês, et al.
Veröffentlicht: (2024)
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
von: Papp, Pál András, et al.
Veröffentlicht: (2024)
von: Papp, Pál András, et al.
Veröffentlicht: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
von: Böhnlein, Toni, et al.
Veröffentlicht: (2025)
von: Böhnlein, Toni, et al.
Veröffentlicht: (2025)
Multiprocessor Scheduling with Memory Constraints: Fundamental Properties and Finding Optimal Solutions
von: Papp, Pál András, et al.
Veröffentlicht: (2025)
von: Papp, Pál András, et al.
Veröffentlicht: (2025)
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
von: Devarakonda, Aditya, et al.
Veröffentlicht: (2025)
von: Devarakonda, Aditya, et al.
Veröffentlicht: (2025)
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
Practical Livelock Analysis in Parameterized Unidirectional Rings
von: Farahat, Aly
Veröffentlicht: (2026)
von: Farahat, Aly
Veröffentlicht: (2026)
D&A: Resource Optimisation in Personalised PageRank Computations Using Multi-Core Machines
von: Yow, Kai Siong, et al.
Veröffentlicht: (2024)
von: Yow, Kai Siong, et al.
Veröffentlicht: (2024)
Massively Parallel Genetic Optimization through Asynchronous Propagation of Populations
von: Taubert, Oskar, et al.
Veröffentlicht: (2023)
von: Taubert, Oskar, et al.
Veröffentlicht: (2023)
$Δ$-Nets: Interaction-Based System for Optimal Parallel $λ$-Reduction
von: Salvadori, Daniel Augusto Rizzi
Veröffentlicht: (2025)
von: Salvadori, Daniel Augusto Rizzi
Veröffentlicht: (2025)
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
von: Thomaz, Gabriel Fernandes, et al.
Veröffentlicht: (2026)
von: Thomaz, Gabriel Fernandes, et al.
Veröffentlicht: (2026)
DDU-Net: A Domain Decomposition-Based CNN for High-Resolution Image Segmentation on Multiple GPUs
von: Verburg, Corné, et al.
Veröffentlicht: (2024)
von: Verburg, Corné, et al.
Veröffentlicht: (2024)
Dynamic Approximate Maximum Matching in the Distributed Vertex Partition Model
von: Robinson, Peter, et al.
Veröffentlicht: (2025)
von: Robinson, Peter, et al.
Veröffentlicht: (2025)
Distributed Tomographic Reconstruction with Quantization
von: Miao, Runxuan, et al.
Veröffentlicht: (2024)
von: Miao, Runxuan, et al.
Veröffentlicht: (2024)
Serial Parallel Reliability Redundancy Allocation Optimization for Energy Efficient and Fault Tolerant Cloud Computing
von: Krishna, Gutha Jaya
Veröffentlicht: (2024)
von: Krishna, Gutha Jaya
Veröffentlicht: (2024)
Replication in Graph Partitioning and Scheduling Problems
von: Papp, Pál András, et al.
Veröffentlicht: (2026)
von: Papp, Pál András, et al.
Veröffentlicht: (2026)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
In search of the lost tree: Hardness and relaxation of spanning trees in temporal graphs
von: Casteigts, Arnaud, et al.
Veröffentlicht: (2023)
von: Casteigts, Arnaud, et al.
Veröffentlicht: (2023)
Simple, strict, proper, happy: A study of reachability in temporal graphs
von: Casteigts, Arnaud, et al.
Veröffentlicht: (2022)
von: Casteigts, Arnaud, et al.
Veröffentlicht: (2022)
Parallel Self-Avoiding Walks for a Low-Autocorrelation Binary Sequences Problem
von: Bošković, Borko, et al.
Veröffentlicht: (2022)
von: Bošković, Borko, et al.
Veröffentlicht: (2022)
A Simple Communication Scheme for Distributed Fast Multipole Methods
von: Kailasa, Srinath
Veröffentlicht: (2026)
von: Kailasa, Srinath
Veröffentlicht: (2026)
Ähnliche Einträge
-
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
von: Zhao, Haisha, et al.
Veröffentlicht: (2025) -
Real-time chaotic video encryption based on multithreaded parallel confusion and diffusion
von: Jiang, Dong, et al.
Veröffentlicht: (2023) -
GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN
von: Yin, Jia, et al.
Veröffentlicht: (2025) -
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
von: Kriemann, Ronald
Veröffentlicht: (2024) -
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)