Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
Fuente:
arXiv
Saved in:
| Main Authors: | Ansari, Mufakir Qamar, Ansari, Mudabir Qamar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
by: Ansari, Mufakir Qamar, et al.
Published: (2025)
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
by: Thomaz, Gabriel Fernandes, et al.
Published: (2026)
by: Thomaz, Gabriel Fernandes, et al.
Published: (2026)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017)
by: Cicirello, Vincent A.
Published: (2017)
GPU-Initiated Networking for NCCL
by: Hamidouche, Khaled, et al.
Published: (2025)
by: Hamidouche, Khaled, et al.
Published: (2025)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
by: Kriemann, Ronald
Published: (2024)
by: Kriemann, Ronald
Published: (2024)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
by: Shakeri, Heman, et al.
Published: (2026)
by: Shakeri, Heman, et al.
Published: (2026)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
by: Goldman, Amos, et al.
Published: (2026)
by: Goldman, Amos, et al.
Published: (2026)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
by: Chillarón, Mónica, et al.
Published: (2024)
by: Chillarón, Mónica, et al.
Published: (2024)
Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion
by: Lappi, Oskar, et al.
Published: (2025)
by: Lappi, Oskar, et al.
Published: (2025)
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
by: Lendve, Shardul, et al.
Published: (2024)
by: Lendve, Shardul, et al.
Published: (2024)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
by: Suzuki, Kengo, et al.
Published: (2026)
by: Suzuki, Kengo, et al.
Published: (2026)
Algorithms for Parallel Shared-Memory Sparse Matrix-Vector Multiplication on Unstructured Matrices
by: Bergmans, Kobe, et al.
Published: (2025)
by: Bergmans, Kobe, et al.
Published: (2025)
ML-Based Optimum Number of CUDA Streams for the GPU Implementation of the Tridiagonal Partition Method
by: Veneva, Milena, et al.
Published: (2025)
by: Veneva, Milena, et al.
Published: (2025)
ML-Based Optimum Sub-system Size Heuristic for the GPU Implementation of the Tridiagonal Partition Method
by: Veneva, Milena
Published: (2025)
by: Veneva, Milena
Published: (2025)
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
by: Venkat, Sreeram, et al.
Published: (2025)
by: Venkat, Sreeram, et al.
Published: (2025)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
by: Papp, Pál András, et al.
Published: (2024)
by: Papp, Pál András, et al.
Published: (2024)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
A Virtual Processor brings back the Free Lunch
by: Kutschbach, Haymo
Published: (2026)
by: Kutschbach, Haymo
Published: (2026)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Leveraging Multi-Instance GPUs through moldable task scheduling
by: Villarrubia, Jorge, et al.
Published: (2025)
by: Villarrubia, Jorge, et al.
Published: (2025)
CLAIRE: Scalable GPU-Accelerated Algorithms for Diffeomorphic Image Registration in 3D
by: Mang, Andreas
Published: (2024)
by: Mang, Andreas
Published: (2024)
Multistep schemes for solving backward stochastic differential equations on GPU
by: Kapllani, Lorenc, et al.
Published: (2019)
by: Kapllani, Lorenc, et al.
Published: (2019)
Multiprocessor Scheduling with Memory Constraints: Fundamental Properties and Finding Optimal Solutions
by: Papp, Pál András, et al.
Published: (2025)
by: Papp, Pál András, et al.
Published: (2025)
Replication in Graph Partitioning and Scheduling Problems
by: Papp, Pál András, et al.
Published: (2026)
by: Papp, Pál András, et al.
Published: (2026)
GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN
by: Yin, Jia, et al.
Published: (2025)
by: Yin, Jia, et al.
Published: (2025)
Population Protocols Revisited: Parity and Beyond
by: Gąsieniec, Leszek, et al.
Published: (2025)
by: Gąsieniec, Leszek, et al.
Published: (2025)
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
by: Ferreira, Inês, et al.
Published: (2024)
by: Ferreira, Inês, et al.
Published: (2024)
Distributed Tomographic Reconstruction with Quantization
by: Miao, Runxuan, et al.
Published: (2024)
by: Miao, Runxuan, et al.
Published: (2024)
CompressedScaffnew: The First Theoretical Double Acceleration of Communication from Local Training and Compression in Distributed Optimization
by: Condat, Laurent, et al.
Published: (2022)
by: Condat, Laurent, et al.
Published: (2022)
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
by: Devarakonda, Aditya, et al.
Published: (2025)
by: Devarakonda, Aditya, et al.
Published: (2025)
Distributed Hybrid Sketching for $\ell_2$-Embeddings
by: Charalambides, Neophytos, et al.
Published: (2024)
by: Charalambides, Neophytos, et al.
Published: (2024)
cuGenOpt: A GPU-Accelerated General-Purpose Metaheuristic Framework for Combinatorial Optimization
by: Liu, Yuyang
Published: (2026)
by: Liu, Yuyang
Published: (2026)
Parallel Self-Avoiding Walks for a Low-Autocorrelation Binary Sequences Problem
by: Bošković, Borko, et al.
Published: (2022)
by: Bošković, Borko, et al.
Published: (2022)
Light Cone Consistency: Toward a Unified Theory of Consistency in Message-Passing Systems
by: Landers, Rob, et al.
Published: (2026)
by: Landers, Rob, et al.
Published: (2026)
Exploring the Design Space for Message-Driven Systems for Dynamic Graph Processing using CCA
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
by: Chandio, Bibrak Qamar, et al.
Published: (2024)
Similar Items
-
Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
by: Ansari, Mufakir Qamar, et al.
Published: (2025) -
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
by: Thomaz, Gabriel Fernandes, et al.
Published: (2026) -
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017) -
GPU-Initiated Networking for NCCL
by: Hamidouche, Khaled, et al.
Published: (2025) -
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
by: Kriemann, Ronald
Published: (2024)