Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach
Fuente:
arXiv
Saved in:
| Main Authors: | Xiaoteng, Liu, Halim, Pavly |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion
by: Lappi, Oskar, et al.
Published: (2025)
by: Lappi, Oskar, et al.
Published: (2025)
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017)
by: Cicirello, Vincent A.
Published: (2017)
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
by: Suzuki, Kengo, et al.
Published: (2026)
by: Suzuki, Kengo, et al.
Published: (2026)
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025)
by: Böhnlein, Toni, et al.
Published: (2025)
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Solving Large Rank-Deficient Linear Least-Squares Problems on Shared-Memory CPU Architectures and GPU Architectures
by: Chillarón, Mónica, et al.
Published: (2024)
by: Chillarón, Mónica, et al.
Published: (2024)
Precision-Aware Iterative Algorithms Based on Group-Shared Exponents of Floating-Point Numbers
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Cascaded Prediction and Asynchronous Execution of Iterative Algorithms on Heterogeneous Platforms
by: Gao, Jianhua, et al.
Published: (2024)
by: Gao, Jianhua, et al.
Published: (2024)
Federated Learning with MMD-based Early Stopping for Adaptive GNSS Interference Classification
by: Gaikwad, Nishant S., et al.
Published: (2024)
by: Gaikwad, Nishant S., et al.
Published: (2024)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
by: Kriemann, Ronald
Published: (2024)
by: Kriemann, Ronald
Published: (2024)
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization
by: Devarakonda, Aditya, et al.
Published: (2025)
by: Devarakonda, Aditya, et al.
Published: (2025)
A Simple Communication Scheme for Distributed Fast Multipole Methods
by: Kailasa, Srinath
Published: (2026)
by: Kailasa, Srinath
Published: (2026)
Flexible Multi-Dimensional FFTs for Plane Wave Density Functional Theory Codes
by: Popovici, Doru Thom, et al.
Published: (2024)
by: Popovici, Doru Thom, et al.
Published: (2024)
An Incrementally Expanding Approach for Updating PageRank on Dynamic Graphs
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
DF* PageRank: Improved Incrementally Expanding Approaches for Updating PageRank on Dynamic Graphs
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
Vectorized Adaptive Histograms for Sparse Oblique Forests
by: Lubonja, Ariel, et al.
Published: (2026)
by: Lubonja, Ariel, et al.
Published: (2026)
Distributed Tomographic Reconstruction with Quantization
by: Miao, Runxuan, et al.
Published: (2024)
by: Miao, Runxuan, et al.
Published: (2024)
DDU-Net: A Domain Decomposition-Based CNN for High-Resolution Image Segmentation on Multiple GPUs
by: Verburg, Corné, et al.
Published: (2024)
by: Verburg, Corné, et al.
Published: (2024)
GVE-LPA: Fast Label Propagation Algorithm (LPA) for Community Detection in Shared Memory Setting
by: Sahu, Subhajit
Published: (2023)
by: Sahu, Subhajit
Published: (2023)
GVE-Louvain: Fast Louvain Algorithm for Community Detection in Shared Memory Setting
by: Sahu, Subhajit
Published: (2023)
by: Sahu, Subhajit
Published: (2023)
GVE-Leiden: Fast Leiden Algorithm for Community Detection in Shared Memory Setting
by: Sahu, Subhajit
Published: (2023)
by: Sahu, Subhajit
Published: (2023)
Lock-Free Computation of PageRank in Dynamic Graphs
by: Sahu, Subhajit
Published: (2024)
by: Sahu, Subhajit
Published: (2024)
Gradient Coding with Iterative Block Leverage Score Sampling
by: Charalambides, Neophytos, et al.
Published: (2023)
by: Charalambides, Neophytos, et al.
Published: (2023)
Accelerated Spatio-Temporal Bayesian Modeling for Multivariate Gaussian Processes
by: Gaedke-Merzhäuser, Lisa, et al.
Published: (2025)
by: Gaedke-Merzhäuser, Lisa, et al.
Published: (2025)
Distributed Hybrid Sketching for $\ell_2$-Embeddings
by: Charalambides, Neophytos, et al.
Published: (2024)
by: Charalambides, Neophytos, et al.
Published: (2024)
Dynamic Memory Management on GPUs with SYCL
by: Standish, Russell K.
Published: (2025)
by: Standish, Russell K.
Published: (2025)
Mixed-Precision Performance Portability of FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices
by: Venkat, Sreeram, et al.
Published: (2025)
by: Venkat, Sreeram, et al.
Published: (2025)
A Hybrid Direct-Iterative Method for Solving KKT Linear Systems
by: Regev, Shaked, et al.
Published: (2021)
by: Regev, Shaked, et al.
Published: (2021)
Serial Parallel Reliability Redundancy Allocation Optimization for Energy Efficient and Fault Tolerant Cloud Computing
by: Krishna, Gutha Jaya
Published: (2024)
by: Krishna, Gutha Jaya
Published: (2024)
A Comparative Analysis of Distributed Linear Solvers under Data Heterogeneity
by: Velasevic, Boris, et al.
Published: (2023)
by: Velasevic, Boris, et al.
Published: (2023)
Population Protocols Revisited: Parity and Beyond
by: Gąsieniec, Leszek, et al.
Published: (2025)
by: Gąsieniec, Leszek, et al.
Published: (2025)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
by: Shakeri, Heman, et al.
Published: (2026)
by: Shakeri, Heman, et al.
Published: (2026)
Parallel Self-Avoiding Walks for a Low-Autocorrelation Binary Sequences Problem
by: Bošković, Borko, et al.
Published: (2022)
by: Bošković, Borko, et al.
Published: (2022)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
by: Cheng, Long, et al.
Published: (2026)
by: Cheng, Long, et al.
Published: (2026)
GPU acceleration of non-equilibrium Green's function calculation using OpenACC and CUDA FORTRAN
by: Yin, Jia, et al.
Published: (2025)
by: Yin, Jia, et al.
Published: (2025)
OPTIMUM-DERAM: Highly Consistent, Scalable, and Secure Multi-Object Memory using RLNC
by: Nicolaou, Nicolas, et al.
Published: (2026)
by: Nicolaou, Nicolas, et al.
Published: (2026)
Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor Cores
by: Zhao, Haisha, et al.
Published: (2025)
by: Zhao, Haisha, et al.
Published: (2025)
Context Adaptive Cooperation
by: Albouy, Timothé, et al.
Published: (2023)
by: Albouy, Timothé, et al.
Published: (2023)
Reinforcement Learning Controlled Adaptive PSO for Task Offloading in IIoT Edge Computing
by: Perera, Minod, et al.
Published: (2025)
by: Perera, Minod, et al.
Published: (2025)
Direct Low-Dose CT Image Reconstruction on GPU using Out-Of-Core: Precision and Quality Study
by: Chillarón, M., et al.
Published: (2024)
by: Chillarón, M., et al.
Published: (2024)
Similar Items
-
Scalable Domain-decomposed Monte Carlo Neutral Transport for Nuclear Fusion
by: Lappi, Oskar, et al.
Published: (2025) -
Design, Configuration, Implementation, and Performance of a Simple 32 Core Raspberry Pi Cluster
by: Cicirello, Vincent A.
Published: (2017) -
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
by: Suzuki, Kengo, et al.
Published: (2026) -
Efficient Parallel Scheduling for Sparse Triangular Solvers
by: Böhnlein, Toni, et al.
Published: (2025) -
A Systematic Literature Survey of Sparse Matrix-Vector Multiplication
by: Gao, Jianhua, et al.
Published: (2024)