Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | Lurati, Milo, Heldens, Stijn, Sclocco, Alessio, van Werkhoven, Ben |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025)
by: Willemsen, Floris-Jan, et al.
Published: (2025)
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
by: Lin, Wei-Chen, et al.
Published: (2024)
by: Lin, Wei-Chen, et al.
Published: (2024)
Efficient Construction of Large Search Spaces for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025)
by: Willemsen, Floris-Jan, et al.
Published: (2025)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
by: León-Vega, Luis G., et al.
Published: (2024)
by: León-Vega, Luis G., et al.
Published: (2024)
How to Rent GPUs on a Budget
by: Li, Zhouzi, et al.
Published: (2024)
by: Li, Zhouzi, et al.
Published: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
by: Wahlgren, Jacob, et al.
Published: (2025)
by: Wahlgren, Jacob, et al.
Published: (2025)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
by: Afzal, Ayesha, et al.
Published: (2025)
by: Afzal, Ayesha, et al.
Published: (2025)
DIAL: Decentralized I/O AutoTuning via Learned Client-side Local Metrics for Parallel File System
by: Rashid, Md Hasanur, et al.
Published: (2026)
by: Rashid, Md Hasanur, et al.
Published: (2026)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
by: Rose, Martin, et al.
Published: (2025)
by: Rose, Martin, et al.
Published: (2025)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
by: Owen, Herbert, et al.
Published: (2024)
by: Owen, Herbert, et al.
Published: (2024)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
by: Katagiri, Takahiro, et al.
Published: (2024)
by: Katagiri, Takahiro, et al.
Published: (2024)
Inductive Loop Analysis for Practical HPC Application Optimization
by: Schaad, Philipp, et al.
Published: (2025)
by: Schaad, Philipp, et al.
Published: (2025)
Seamless acceleration of Fortran intrinsics via AMD AI engines
by: Brown, Nick, et al.
Published: (2025)
by: Brown, Nick, et al.
Published: (2025)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
by: Anik, Md Saidul Hoque, et al.
Published: (2024)
by: Anik, Md Saidul Hoque, et al.
Published: (2024)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
by: Chen, David, et al.
Published: (2026)
by: Chen, David, et al.
Published: (2026)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
by: Rashid, Md Hasanur, et al.
Published: (2026)
by: Rashid, Md Hasanur, et al.
Published: (2026)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
by: Qararyah, Fareed, et al.
Published: (2024)
by: Qararyah, Fareed, et al.
Published: (2024)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
by: Peng, Hongwu, et al.
Published: (2023)
by: Peng, Hongwu, et al.
Published: (2023)
Optimizing sDTW for AMD GPUs
by: Latta-Lin, Daniel, et al.
Published: (2024)
by: Latta-Lin, Daniel, et al.
Published: (2024)
Performance Impact of Containerized METADOCK 2 on Heterogeneous Platforms
by: Banegas-Luna, Antonio Jesús, et al.
Published: (2025)
by: Banegas-Luna, Antonio Jesús, et al.
Published: (2025)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
by: Zhu, Jianwei, et al.
Published: (2024)
by: Zhu, Jianwei, et al.
Published: (2024)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
by: Papavasileiou, Ioannis, et al.
Published: (2026)
by: Papavasileiou, Ioannis, et al.
Published: (2026)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
by: Panova, Elena, et al.
Published: (2022)
by: Panova, Elena, et al.
Published: (2022)
Staging Blocked Evaluation over Structured Sparse Matrices
by: Das, Pratyush, et al.
Published: (2024)
by: Das, Pratyush, et al.
Published: (2024)
A Performance Analysis of BFT Consensus for Blockchains
by: Chan, J. D., et al.
Published: (2024)
by: Chan, J. D., et al.
Published: (2024)
Scalable GPU Performance Variability Analysis framework
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
Automated Programmatic Performance Analysis of Parallel Programs
by: Cankur, Onur, et al.
Published: (2024)
by: Cankur, Onur, et al.
Published: (2024)
Performance Debugging through Microarchitectural Sensitivity and Causality Analysis
by: Dutilleul, Alban, et al.
Published: (2024)
by: Dutilleul, Alban, et al.
Published: (2024)
Recorder: Comprehensive Parallel I/O Tracing and Analysis
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
by: Lin, Mao, et al.
Published: (2026)
by: Lin, Mao, et al.
Published: (2026)
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
by: Battaglini-Fischer, Sándor, et al.
Published: (2025)
by: Battaglini-Fischer, Sándor, et al.
Published: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
by: Islam, Tanzima Z., et al.
Published: (2024)
by: Islam, Tanzima Z., et al.
Published: (2024)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
by: Afzal, Ayesha, et al.
Published: (2024)
by: Afzal, Ayesha, et al.
Published: (2024)
Tuning of Vectorization Parameters for Molecular Dynamics Simulations in AutoPas
by: Gall, Luis, et al.
Published: (2025)
by: Gall, Luis, et al.
Published: (2025)
Performance and scaling of the LFRic weather and climate model on different generations of HPE Cray EX supercomputers
by: Bull, J. Mark, et al.
Published: (2024)
by: Bull, J. Mark, et al.
Published: (2024)
Constructive community race: full-density spiking neural network model drives neuromorphic computing
by: Senk, Johanna, et al.
Published: (2025)
by: Senk, Johanna, et al.
Published: (2025)
AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference
by: Zhao, Xuanlei, et al.
Published: (2024)
by: Zhao, Xuanlei, et al.
Published: (2024)
DiFuseR: A Distributed Sketch-based Influence Maximization Algorithm for GPUs
by: Göktürk, Gökhan, et al.
Published: (2024)
by: Göktürk, Gökhan, et al.
Published: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
by: Gupta, Ahan, et al.
Published: (2026)
by: Gupta, Ahan, et al.
Published: (2026)
Similar Items
-
Tuning the Tuner: Introducing Hyperparameter Optimization for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025) -
Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
by: Lin, Wei-Chen, et al.
Published: (2024) -
Efficient Construction of Large Search Spaces for Auto-Tuning
by: Willemsen, Floris-Jan, et al.
Published: (2025) -
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
by: Katagiri, Takahiro, et al.
Published: (2024) -
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
by: León-Vega, Luis G., et al.
Published: (2024)