Preliminary report: Initial evaluation of StdPar implementations on AMD GPUs for HPC
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lin, Wei-Chen, McIntosh-Smith, Simon, Deakin, Tom |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
par: Lurati, Milo, et autres
Publié: (2024)
par: Lurati, Milo, et autres
Publié: (2024)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
par: Rose, Martin, et autres
Publié: (2025)
par: Rose, Martin, et autres
Publié: (2025)
Usability Evaluation of Cloud for HPC Applications
par: Sochat, Vanessa, et autres
Publié: (2025)
par: Sochat, Vanessa, et autres
Publié: (2025)
How to Rent GPUs on a Budget
par: Li, Zhouzi, et autres
Publié: (2024)
par: Li, Zhouzi, et autres
Publié: (2024)
Extrae.jl: Julia bindings for the Extrae HPC Profiler
par: Sanchez-Ramirez, Sergio, et autres
Publié: (2025)
par: Sanchez-Ramirez, Sergio, et autres
Publié: (2025)
Inductive Loop Analysis for Practical HPC Application Optimization
par: Schaad, Philipp, et autres
Publié: (2025)
par: Schaad, Philipp, et autres
Publié: (2025)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
par: Wahlgren, Jacob, et autres
Publié: (2025)
par: Wahlgren, Jacob, et autres
Publié: (2025)
LLload: Simplifying Real-Time Job Monitoring for HPC Users
par: Byun, Chansup, et autres
Publié: (2024)
par: Byun, Chansup, et autres
Publié: (2024)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
par: Tharwani, Jay, et autres
Publié: (2025)
par: Tharwani, Jay, et autres
Publié: (2025)
Operational Strategies for Non-Disruptive Scheduling Transitions in Production HPC Systems
par: MacLachlan, Glen, et autres
Publié: (2026)
par: MacLachlan, Glen, et autres
Publié: (2026)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
par: Afzal, Ayesha, et autres
Publié: (2025)
par: Afzal, Ayesha, et autres
Publié: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
par: Ather, Hammad, et autres
Publié: (2024)
par: Ather, Hammad, et autres
Publié: (2024)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
par: Jain, Rutwik, et autres
Publié: (2026)
par: Jain, Rutwik, et autres
Publié: (2026)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
par: León-Vega, Luis G., et autres
Publié: (2024)
par: León-Vega, Luis G., et autres
Publié: (2024)
Alya towards Exascale: Optimal OpenACC Performance of the Navier-Stokes Finite Element Assembly on GPUs
par: Owen, Herbert, et autres
Publié: (2024)
par: Owen, Herbert, et autres
Publié: (2024)
Optimizing sDTW for AMD GPUs
par: Latta-Lin, Daniel, et autres
Publié: (2024)
par: Latta-Lin, Daniel, et autres
Publié: (2024)
Seamless acceleration of Fortran intrinsics via AMD AI engines
par: Brown, Nick, et autres
Publié: (2025)
par: Brown, Nick, et autres
Publié: (2025)
An experimental evaluation of satellite constellation emulators
par: Cionca, Victor, et autres
Publié: (2026)
par: Cionca, Victor, et autres
Publié: (2026)
GROMACS on AMD GPU-Based HPC Platforms: Using SYCL for Performance and Portability
par: Alekseenko, Andrey, et autres
Publié: (2024)
par: Alekseenko, Andrey, et autres
Publié: (2024)
Evaluating Emerging AI/ML Accelerators: IPU, RDU, and NVIDIA/AMD GPUs
par: Peng, Hongwu, et autres
Publié: (2023)
par: Peng, Hongwu, et autres
Publié: (2023)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
par: Qararyah, Fareed, et autres
Publié: (2024)
par: Qararyah, Fareed, et autres
Publié: (2024)
Confidential Computing on NVIDIA Hopper GPUs: A Performance Benchmark Study
par: Zhu, Jianwei, et autres
Publié: (2024)
par: Zhu, Jianwei, et autres
Publié: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
par: Nicusan, Andrei-Leonard, et autres
Publié: (2025)
par: Nicusan, Andrei-Leonard, et autres
Publié: (2025)
AdapTBF: Decentralized Bandwidth Control via Adaptive Token Borrowing for HPC Storage
par: Rashid, Md Hasanur, et autres
Publié: (2026)
par: Rashid, Md Hasanur, et autres
Publié: (2026)
Black-Scholes Option Pricing on Intel CPUs and GPUs: Implementation on SYCL and Optimization Techniques
par: Panova, Elena, et autres
Publié: (2022)
par: Panova, Elena, et autres
Publié: (2022)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
par: Oztop, Beste, et autres
Publié: (2026)
par: Oztop, Beste, et autres
Publié: (2026)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
par: Ibeid, Huda, et autres
Publié: (2025)
par: Ibeid, Huda, et autres
Publié: (2025)
PASTA: A Modular Program Analysis Tool Framework for Accelerators
par: Lin, Mao, et autres
Publié: (2026)
par: Lin, Mao, et autres
Publié: (2026)
Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study
par: McDonald, Jesse, et autres
Publié: (2024)
par: McDonald, Jesse, et autres
Publié: (2024)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
par: Lin, Mao, et autres
Publié: (2026)
par: Lin, Mao, et autres
Publié: (2026)
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study
par: Debnath, Shimul, et autres
Publié: (2026)
par: Debnath, Shimul, et autres
Publié: (2026)
Isambard-AI: a leadership class supercomputer optimised specifically for Artificial Intelligence
par: McIntosh-Smith, Simon, et autres
Publié: (2024)
par: McIntosh-Smith, Simon, et autres
Publié: (2024)
Federated Single Sign-On and Zero Trust Co-design for AI and HPC Digital Research Infrastructures
par: Alam, Sadaf R., et autres
Publié: (2024)
par: Alam, Sadaf R., et autres
Publié: (2024)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
par: Zhuang, Chen, et autres
Publié: (2025)
par: Zhuang, Chen, et autres
Publié: (2025)
On Orchestrating Parallel Broadcasts for Distributed Ledgers
par: Sheng, Peiyao, et autres
Publié: (2024)
par: Sheng, Peiyao, et autres
Publié: (2024)
Constructive community race: full-density spiking neural network model drives neuromorphic computing
par: Senk, Johanna, et autres
Publié: (2025)
par: Senk, Johanna, et autres
Publié: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
par: Zhang, Li, et autres
Publié: (2025)
par: Zhang, Li, et autres
Publié: (2025)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
par: Zhang, Yaozheng, et autres
Publié: (2025)
par: Zhang, Yaozheng, et autres
Publié: (2025)
Performance Optimization in Stream Processing Systems: Experiment-Driven Configuration Tuning for Kafka Streams
par: Chen, David, et autres
Publié: (2026)
par: Chen, David, et autres
Publié: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
par: Zhuang, Chen, et autres
Publié: (2024)
par: Zhuang, Chen, et autres
Publié: (2024)
Documents similaires
-
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
par: Lurati, Milo, et autres
Publié: (2024) -
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
par: Rose, Martin, et autres
Publié: (2025) -
Usability Evaluation of Cloud for HPC Applications
par: Sochat, Vanessa, et autres
Publié: (2025) -
How to Rent GPUs on a Budget
par: Li, Zhouzi, et autres
Publié: (2024) -
Extrae.jl: Julia bindings for the Extrae HPC Profiler
par: Sanchez-Ramirez, Sergio, et autres
Publié: (2025)