Redesigning GROMACS Halo Exchange: Improving Strong Scaling with GPU-initiated NVSHMEM
Fuente:
arXiv
Saved in:
| Main Authors: | Doijade, Mahesh, Alekseenko, Andrey, Brown, Ania, Gray, Alan, Páll, Szilárd |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GROMACS on AMD GPU-Based HPC Platforms: Using SYCL for Performance and Portability
by: Alekseenko, Andrey, et al.
Published: (2024)
by: Alekseenko, Andrey, et al.
Published: (2024)
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
by: Afzal, Ayesha, et al.
Published: (2025)
by: Afzal, Ayesha, et al.
Published: (2025)
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
by: Apanasevich, L., et al.
Published: (2024)
by: Apanasevich, L., et al.
Published: (2024)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
by: Villalobos, Johansell, et al.
Published: (2025)
by: Villalobos, Johansell, et al.
Published: (2025)
The Energy Cost of Execution-Idle in GPU Clusters
by: Lei, Yiran, et al.
Published: (2026)
by: Lei, Yiran, et al.
Published: (2026)
Scalable GPU Performance Variability Analysis framework
by: Lahiry, Ankur, et al.
Published: (2025)
by: Lahiry, Ankur, et al.
Published: (2025)
On the Partitioning of GPU Power among Multi-Instances
by: Vamja, Tirth, et al.
Published: (2025)
by: Vamja, Tirth, et al.
Published: (2025)
Disaggregated Design for GPU-Based Volumetric Data Structures
by: Meneghin, Massimiliano, et al.
Published: (2025)
by: Meneghin, Massimiliano, et al.
Published: (2025)
Taking GPU Programming Models to Task for Performance Portability
by: Davis, Joshua H., et al.
Published: (2024)
by: Davis, Joshua H., et al.
Published: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
by: Lawenda, Marcin, et al.
Published: (2025)
by: Lawenda, Marcin, et al.
Published: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
by: Zhao, Yanbo, et al.
Published: (2025)
by: Zhao, Yanbo, et al.
Published: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
by: Davis, Joshua H., et al.
Published: (2026)
by: Davis, Joshua H., et al.
Published: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
by: Pilliat, Emmanuel
Published: (2026)
by: Pilliat, Emmanuel
Published: (2026)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
by: Lawenda, Marcin, et al.
Published: (2025)
by: Lawenda, Marcin, et al.
Published: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
by: Islam, Tanzima Z., et al.
Published: (2024)
by: Islam, Tanzima Z., et al.
Published: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
by: Liu, Shifang, et al.
Published: (2025)
by: Liu, Shifang, et al.
Published: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
by: Wang, Yidi, et al.
Published: (2024)
by: Wang, Yidi, et al.
Published: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
by: Wahlgren, Jacob, et al.
Published: (2025)
by: Wahlgren, Jacob, et al.
Published: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
by: Xia, Yuning, et al.
Published: (2026)
by: Xia, Yuning, et al.
Published: (2026)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
by: Curless, Brian, et al.
Published: (2025)
by: Curless, Brian, et al.
Published: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
by: Lin, Mao, et al.
Published: (2026)
by: Lin, Mao, et al.
Published: (2026)
Vectorization of Gradient Boosting of Decision Trees Prediction in the CatBoost Library for RISC-V Processors
by: Kozinov, Evgeny, et al.
Published: (2024)
by: Kozinov, Evgeny, et al.
Published: (2024)
RAID Organizations for Improved Reliability and Performance: A Not Entirely Unbiased Tutorial
by: Thomasian, Alexander
Published: (2023)
by: Thomasian, Alexander
Published: (2023)
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
by: Vatsavai, Sairam Sri, et al.
Published: (2025)
by: Vatsavai, Sairam Sri, et al.
Published: (2025)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
by: Arif, Moiz, et al.
Published: (2026)
by: Arif, Moiz, et al.
Published: (2026)
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale Diagnostics
by: Grbic, Dragana
Published: (2026)
by: Grbic, Dragana
Published: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
by: Zhuang, Chen, et al.
Published: (2024)
by: Zhuang, Chen, et al.
Published: (2024)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023)
by: Wang, Yuxin, et al.
Published: (2023)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
by: Xu, Jingwei, et al.
Published: (2025)
by: Xu, Jingwei, et al.
Published: (2025)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
by: Suffa, Philipp, et al.
Published: (2024)
by: Suffa, Philipp, et al.
Published: (2024)
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
by: Yildiz, Mert, et al.
Published: (2025)
by: Yildiz, Mert, et al.
Published: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
by: Ather, Hammad, et al.
Published: (2024)
by: Ather, Hammad, et al.
Published: (2024)
Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
by: Mandarapu, Durga, et al.
Published: (2023)
by: Mandarapu, Durga, et al.
Published: (2023)
A Precision Emulation Approach to the GPU Acceleration of Ab Initio Electronic Structure Calculations
by: Liu, Hang, et al.
Published: (2026)
by: Liu, Hang, et al.
Published: (2026)
The Landscape of GPU-Centric Communication
by: Unat, Didem, et al.
Published: (2024)
by: Unat, Didem, et al.
Published: (2024)
GigaAPI for GPU Parallelization
by: Suvarna, M., et al.
Published: (2025)
by: Suvarna, M., et al.
Published: (2025)
Optimizing Near Field Computation in the MLFMA Algorithm with Data Redundancy and Performance Modeling on a Single GPU
by: Sadeghi, Morteza, et al.
Published: (2024)
by: Sadeghi, Morteza, et al.
Published: (2024)
Parallelizing a modern GPU simulator
by: Huerta, Rodrigo, et al.
Published: (2025)
by: Huerta, Rodrigo, et al.
Published: (2025)
Accelerating Particle-in-Cell Monte Carlo Simulations with MPI, OpenMP/OpenACC and Asynchronous Multi-GPU Programming
by: Williams, Jeremy J., et al.
Published: (2024)
by: Williams, Jeremy J., et al.
Published: (2024)
Similar Items
-
GROMACS on AMD GPU-Based HPC Platforms: Using SYCL for Performance and Portability
by: Alekseenko, Andrey, et al.
Published: (2024) -
GROMACS Unplugged: How Power Capping and Frequency Shapes Performance on GPUs
by: Afzal, Ayesha, et al.
Published: (2025) -
A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
by: Apanasevich, L., et al.
Published: (2024) -
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
by: Villalobos, Johansell, et al.
Published: (2025) -
The Energy Cost of Execution-Idle in GPU Clusters
by: Lei, Yiran, et al.
Published: (2026)