Racing to Idle: Energy Efficiency of Matrix Multiplication on Heterogeneous CPU and GPU Architectures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ansari, Mufakir Qamar, Ansari, Mudabir Qamar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
von: Kriemann, Ronald
Veröffentlicht: (2024)
von: Kriemann, Ronald
Veröffentlicht: (2024)
Algorithms for Parallel Shared-Memory Sparse Matrix-Vector Multiplication on Unstructured Matrices
von: Bergmans, Kobe, et al.
Veröffentlicht: (2025)
von: Bergmans, Kobe, et al.
Veröffentlicht: (2025)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
von: Zhao, Haisha, et al.
Veröffentlicht: (2025)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN
von: Yin, Jia, et al.
Veröffentlicht: (2025)
von: Yin, Jia, et al.
Veröffentlicht: (2025)
Population Protocols Revisited: Parity and Beyond
von: Gąsieniec, Leszek, et al.
Veröffentlicht: (2025)
von: Gąsieniec, Leszek, et al.
Veröffentlicht: (2025)
Distributed Tomographic Reconstruction with Quantization
von: Miao, Runxuan, et al.
Veröffentlicht: (2024)
von: Miao, Runxuan, et al.
Veröffentlicht: (2024)
Dynamic Memory Management on GPUs with SYCL
von: Standish, Russell K.
Veröffentlicht: (2025)
von: Standish, Russell K.
Veröffentlicht: (2025)
ML-Based Optimum Number of CUDA Streams for the GPU Implementation of the Tridiagonal Partition Method
von: Veneva, Milena, et al.
Veröffentlicht: (2025)
von: Veneva, Milena, et al.
Veröffentlicht: (2025)
ML-Based Optimum Sub-system Size Heuristic for the GPU Implementation of the Tridiagonal Partition Method
von: Veneva, Milena
Veröffentlicht: (2025)
von: Veneva, Milena
Veröffentlicht: (2025)
CLAIRE: Scalable GPU-Accelerated Algorithms for Diffeomorphic Image Registration in 3D
von: Mang, Andreas
Veröffentlicht: (2024)
von: Mang, Andreas
Veröffentlicht: (2024)
Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion
von: Lappi, Oskar, et al.
Veröffentlicht: (2025)
von: Lappi, Oskar, et al.
Veröffentlicht: (2025)
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
von: Ferreira, Inês, et al.
Veröffentlicht: (2024)
von: Ferreira, Inês, et al.
Veröffentlicht: (2024)
OPTIMUM-DERAM: Highly Consistent, Scalable, and Secure Multi-Object Memory using RLNC
von: Nicolaou, Nicolas, et al.
Veröffentlicht: (2026)
von: Nicolaou, Nicolas, et al.
Veröffentlicht: (2026)
Context Adaptive Cooperation
von: Albouy, Timothé, et al.
Veröffentlicht: (2023)
von: Albouy, Timothé, et al.
Veröffentlicht: (2023)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
von: Chillarón, Mónica, et al.
Veröffentlicht: (2024)
von: Chillarón, Mónica, et al.
Veröffentlicht: (2024)
Efficient Parallel Scheduling for Sparse Triangular Solvers
von: Böhnlein, Toni, et al.
Veröffentlicht: (2025)
von: Böhnlein, Toni, et al.
Veröffentlicht: (2025)
A Morton-Type Space-Filling Curve for Pyramid Subdivision and Hybrid Adaptive Mesh Refinement
von: Knapp, David, et al.
Veröffentlicht: (2026)
von: Knapp, David, et al.
Veröffentlicht: (2026)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
von: Venkat, Sreeram, et al.
Veröffentlicht: (2025)
von: Venkat, Sreeram, et al.
Veröffentlicht: (2025)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
von: Cicirello, Vincent A.
Veröffentlicht: (2017)
von: Cicirello, Vincent A.
Veröffentlicht: (2017)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)
Real-time chaotic video encryption based on multithreaded parallel confusion and diffusion
von: Jiang, Dong, et al.
Veröffentlicht: (2023)
von: Jiang, Dong, et al.
Veröffentlicht: (2023)
Massively Parallel Reductions in Multivariate Polynomial Systems: Bridging the Symbolic Preprocessing Gap on GPGPU Architectures
von: Gokavarapu, Chandrasekhar
Veröffentlicht: (2026)
von: Gokavarapu, Chandrasekhar
Veröffentlicht: (2026)
Accelerating State-Vector Quantum Simulation on Integrated GPUs via Cache Locality Optimization: A Cross-Architecture Evaluation
von: Thomaz, Gabriel Fernandes, et al.
Veröffentlicht: (2026)
von: Thomaz, Gabriel Fernandes, et al.
Veröffentlicht: (2026)
Multistep schemes for solving backward stochastic differential equations on GPU
von: Kapllani, Lorenc, et al.
Veröffentlicht: (2019)
von: Kapllani, Lorenc, et al.
Veröffentlicht: (2019)
Dynamic Approximate Maximum Matching in the Distributed Vertex Partition Model
von: Robinson, Peter, et al.
Veröffentlicht: (2025)
von: Robinson, Peter, et al.
Veröffentlicht: (2025)
Data Scheduling Algorithm for Scalable and Efficient IoT Sensing in Cloud Computing
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
GPU-Initiated Networking for NCCL
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
DDU-Net: A Domain Decomposition-Based CNN for High-Resolution Image Segmentation on Multiple GPUs
von: Verburg, Corné, et al.
Veröffentlicht: (2024)
von: Verburg, Corné, et al.
Veröffentlicht: (2024)
A Lock-Free, Fully GPU-Resident Architecture for the Verification of Goldbach's Conjecture
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
von: Suzuki, Kengo, et al.
Veröffentlicht: (2026)
von: Suzuki, Kengo, et al.
Veröffentlicht: (2026)
Massively Parallel Genetic Optimization through Asynchronous Propagation of Populations
von: Taubert, Oskar, et al.
Veröffentlicht: (2023)
von: Taubert, Oskar, et al.
Veröffentlicht: (2023)
Exploiting nested task-parallelism in the $\mathcal{H}-LU$ factorization
von: Carratalá-Sáez, Rocío, et al.
Veröffentlicht: (2019)
von: Carratalá-Sáez, Rocío, et al.
Veröffentlicht: (2019)
Machine Learning-Driven Predictive Resource Management in Complex Science Workflows
von: Chowdhury, Tasnuva, et al.
Veröffentlicht: (2025)
von: Chowdhury, Tasnuva, et al.
Veröffentlicht: (2025)
Stochastic well-structured transition systems
von: Aspnes, James
Veröffentlicht: (2025)
von: Aspnes, James
Veröffentlicht: (2025)
Serial Parallel Reliability Redundancy Allocation Optimization for Energy Efficient and Fault Tolerant Cloud Computing
von: Krishna, Gutha Jaya
Veröffentlicht: (2024)
von: Krishna, Gutha Jaya
Veröffentlicht: (2024)
A Simple Communication Scheme for Distributed Fast Multipole Methods
von: Kailasa, Srinath
Veröffentlicht: (2026)
von: Kailasa, Srinath
Veröffentlicht: (2026)
Ähnliche Einträge
-
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
von: Ansari, Mufakir Qamar, et al.
Veröffentlicht: (2025) -
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
von: Kriemann, Ronald
Veröffentlicht: (2024) -
Algorithms for Parallel Shared-Memory Sparse Matrix-Vector Multiplication on Unstructured Matrices
von: Bergmans, Kobe, et al.
Veröffentlicht: (2025) -
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
von: Zhao, Haisha, et al.
Veröffentlicht: (2025) -
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
von: Gao, Jianhua, et al.
Veröffentlicht: (2024)