Understanding GPU Triggering APIs for MPI+X Communication
Fuente:
arXiv
Saved in:
| Main Authors: | Bridges, Patrick G., Skjellum, Anthony, Suggs, Evan D., Schafer, Derek, Bangalore, Purushotham V. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
by: Bridges, Patrick G., et al.
Published: (2026)
by: Bridges, Patrick G., et al.
Published: (2026)
Concepts for designing modern C++ interfaces for MPI
by: Avans, C. Nicole, et al.
Published: (2025)
by: Avans, C. Nicole, et al.
Published: (2025)
Leveraging Caliper and Benchpark to Analyze MPI Communication Patterns: Insights from AMG2023, Kripke, and Laghos
by: Nansamba, Grace, et al.
Published: (2025)
by: Nansamba, Grace, et al.
Published: (2025)
The Case for ABI Interoperability in a Fault Tolerant MPI
by: Xu, Yao, et al.
Published: (2025)
by: Xu, Yao, et al.
Published: (2025)
MPI Implementation Profiling for Better Application Performance
by: Shipley, Riley, et al.
Published: (2024)
by: Shipley, Riley, et al.
Published: (2024)
A More Scalable Sparse Dynamic Data Exchange
by: Geyko, Andrew, et al.
Published: (2023)
by: Geyko, Andrew, et al.
Published: (2023)
Beatnik: A Novel Global Communication Mini-Application
by: Stewart, Jason A., et al.
Published: (2024)
by: Stewart, Jason A., et al.
Published: (2024)
Persistent and Partitioned MPI for Stencil Communication
by: Collom, Gerald, et al.
Published: (2025)
by: Collom, Gerald, et al.
Published: (2025)
Examining MPI and its Extensions for Asynchronous Multithreaded Communication
by: Yan, Jiakun, et al.
Published: (2025)
by: Yan, Jiakun, et al.
Published: (2025)
Scaling MPI Applications on Aurora
by: Ibeid, Huda, et al.
Published: (2025)
by: Ibeid, Huda, et al.
Published: (2025)
Implementing True MPI Sessions and Evaluating MPI Initialization Scalability
by: Zhou, Hui, et al.
Published: (2026)
by: Zhou, Hui, et al.
Published: (2026)
On the performance of two-sided MPI, MPI-3 RMA and SHMEM in a Lagrangian particle cluster algorithm
by: Frey, Matthias, et al.
Published: (2024)
by: Frey, Matthias, et al.
Published: (2024)
MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
Analyzing Persistent Alltoallv RMA Implementations for High-Performance MPI Communication
by: Namugwanya, Evelyn
Published: (2026)
by: Namugwanya, Evelyn
Published: (2026)
MPI Progress For All
by: Zhou, Hui, et al.
Published: (2024)
by: Zhou, Hui, et al.
Published: (2024)
Communication Round and Computation Efficient Exclusive Prefix-Sums Algorithms (for MPI_Exscan)
by: Träff, Jesper Larsson
Published: (2025)
by: Träff, Jesper Larsson
Published: (2025)
Some New Approaches to MPI Implementations
by: Xiong, Yuqing
Published: (2024)
by: Xiong, Yuqing
Published: (2024)
Designing and Prototyping Extensions to MPI in MPICH
by: Zhou, Hui, et al.
Published: (2024)
by: Zhou, Hui, et al.
Published: (2024)
Frustrated with MPI+Threads? Try MPIxThreads!
by: Zhou, Hui, et al.
Published: (2024)
by: Zhou, Hui, et al.
Published: (2024)
MPI Malleability Validation under Replayed Real-World HPC Conditions
by: Iserte, S., et al.
Published: (2026)
by: Iserte, S., et al.
Published: (2026)
Understanding the Landscape of Ampere GPU Memory Errors
by: Zhu, Zhu, et al.
Published: (2025)
by: Zhu, Zhu, et al.
Published: (2025)
MPI-Q: A Message Communication Library for Large-Scale Classical-Quantum Heterogeneous Hybrid Distributed Computing
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
Parallel Spawning Strategies for Dynamic-Aware MPI Applications
by: Martín-Álvarez, Iker, et al.
Published: (2025)
by: Martín-Álvarez, Iker, et al.
Published: (2025)
Towards the Democratization and Standardization of Dynamic Resources with MPI Spawning
by: Iserte, Sergio, et al.
Published: (2026)
by: Iserte, Sergio, et al.
Published: (2026)
Exploring Performance-Productivity Trade-offs in AMT Runtimes: A Task Bench Study of Itoyori, ItoyoriFBC, HPX, and MPI
by: Lahnor, Torben R., et al.
Published: (2026)
by: Lahnor, Torben R., et al.
Published: (2026)
THAPI: Tracing Heterogeneous APIs
by: Bekele, Solomon, et al.
Published: (2025)
by: Bekele, Solomon, et al.
Published: (2025)
Understanding GPU Resource Interference One Level Deeper
by: Elvinger, Paul, et al.
Published: (2025)
by: Elvinger, Paul, et al.
Published: (2025)
Do MPI Derived Datatypes Actually Help? A Single-Node Cross-Implementation Study on Shared-Memory Communication
by: Adefemi, Temitayo
Published: (2025)
by: Adefemi, Temitayo
Published: (2025)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
To Repair or Not to Repair: Assessing Fault Resilience in MPI Stencil Applications
by: Rocco, Roberto, et al.
Published: (2024)
by: Rocco, Roberto, et al.
Published: (2024)
Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
by: Klepl, Jiří, et al.
Published: (2025)
by: Klepl, Jiří, et al.
Published: (2025)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
by: Song, Boyang
Published: (2025)
by: Song, Boyang
Published: (2025)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
by: Zwaka, Linus
Published: (2024)
by: Zwaka, Linus
Published: (2024)
KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPI
by: Uhl, Tim Niklas, et al.
Published: (2024)
by: Uhl, Tim Niklas, et al.
Published: (2024)
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
by: Iserte, Sergio, et al.
Published: (2025)
by: Iserte, Sergio, et al.
Published: (2025)
Performance of a high-order MPI-Kokkos accelerated fluid solver
by: Sporykhin, Filipp, et al.
Published: (2025)
by: Sporykhin, Filipp, et al.
Published: (2025)
An Optimized Error-controlled MPI Collective Framework Integrated with Lossy Compression
by: Huang, Jiajun, et al.
Published: (2023)
by: Huang, Jiajun, et al.
Published: (2023)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Synthesizing Proxy Applications for MPI Programs
by: Luo, Jiyu, et al.
Published: (2023)
by: Luo, Jiyu, et al.
Published: (2023)
Enhancing Type Safety in MPI with Rust: A Statically Verified Approach for RSMPI
by: Iqbal, Nafees, et al.
Published: (2025)
by: Iqbal, Nafees, et al.
Published: (2025)
Similar Items
-
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
by: Bridges, Patrick G., et al.
Published: (2026) -
Concepts for designing modern C++ interfaces for MPI
by: Avans, C. Nicole, et al.
Published: (2025) -
Leveraging Caliper and Benchpark to Analyze MPI Communication Patterns: Insights from AMG2023, Kripke, and Laghos
by: Nansamba, Grace, et al.
Published: (2025) -
The Case for ABI Interoperability in a Fault Tolerant MPI
by: Xu, Yao, et al.
Published: (2025) -
MPI Implementation Profiling for Better Application Performance
by: Shipley, Riley, et al.
Published: (2024)