Persistent and Partitioned MPI for Stencil Communication
Fuente:
arXiv
Saved in:
| Main Authors: | Collom, Gerald, Burmark, Jason, Pearce, Olga, Bienz, Amanda |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A More Scalable Sparse Dynamic Data Exchange
by: Geyko, Andrew, et al.
Published: (2023)
by: Geyko, Andrew, et al.
Published: (2023)
To Repair or Not to Repair: Assessing Fault Resilience in MPI Stencil Applications
by: Rocco, Roberto, et al.
Published: (2024)
by: Rocco, Roberto, et al.
Published: (2024)
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
by: Adams, Michael, et al.
Published: (2025)
by: Adams, Michael, et al.
Published: (2025)
Analyzing Persistent Alltoallv RMA Implementations for High-Performance MPI Communication
by: Namugwanya, Evelyn
Published: (2026)
by: Namugwanya, Evelyn
Published: (2026)
Stencil Matrixization
by: Zhao, Wenxuan, et al.
Published: (2023)
by: Zhao, Wenxuan, et al.
Published: (2023)
Leveraging Caliper and Benchpark to Analyze MPI Communication Patterns: Insights from AMG2023, Kripke, and Laghos
by: Nansamba, Grace, et al.
Published: (2025)
by: Nansamba, Grace, et al.
Published: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Examining MPI and its Extensions for Asynchronous Multithreaded Communication
by: Yan, Jiakun, et al.
Published: (2025)
by: Yan, Jiakun, et al.
Published: (2025)
Understanding GPU Triggering APIs for MPI+X Communication
by: Bridges, Patrick G., et al.
Published: (2024)
by: Bridges, Patrick G., et al.
Published: (2024)
Do We Need Tensor Cores for Stencil Computations?
by: Gu, Qiqi, et al.
Published: (2026)
by: Gu, Qiqi, et al.
Published: (2026)
Implementing True MPI Sessions and Evaluating MPI Initialization Scalability
by: Zhou, Hui, et al.
Published: (2026)
by: Zhou, Hui, et al.
Published: (2026)
MPI Implementation Profiling for Better Application Performance
by: Shipley, Riley, et al.
Published: (2024)
by: Shipley, Riley, et al.
Published: (2024)
MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
by: Kwon, Miryeong, et al.
Published: (2025)
by: Kwon, Miryeong, et al.
Published: (2025)
MPI Progress For All
by: Zhou, Hui, et al.
Published: (2024)
by: Zhou, Hui, et al.
Published: (2024)
Evaluation of Programming Models and Performance for Stencil Computation on Current GPU Architectures
by: Shan, Baodi, et al.
Published: (2024)
by: Shan, Baodi, et al.
Published: (2024)
A Portable Framework for Accelerating Stencil Computations on Modern Node Architectures
by: Sai, Ryuichi, et al.
Published: (2023)
by: Sai, Ryuichi, et al.
Published: (2023)
Communication Round and Computation Efficient Exclusive Prefix-Sums Algorithms (for MPI_Exscan)
by: Träff, Jesper Larsson
Published: (2025)
by: Träff, Jesper Larsson
Published: (2025)
Scaling MPI Applications on Aurora
by: Ibeid, Huda, et al.
Published: (2025)
by: Ibeid, Huda, et al.
Published: (2025)
MMStencil: Optimizing High-order Stencils on Multicore CPU using Matrix Unit
by: Wang, Yinuo, et al.
Published: (2025)
by: Wang, Yinuo, et al.
Published: (2025)
SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping
by: GU, Qiqi, et al.
Published: (2025)
by: GU, Qiqi, et al.
Published: (2025)
Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
by: Kinkead, Shannon, et al.
Published: (2026)
by: Kinkead, Shannon, et al.
Published: (2026)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
by: Bridges, Patrick G., et al.
Published: (2026)
by: Bridges, Patrick G., et al.
Published: (2026)
On the performance of two-sided MPI, MPI-3 RMA and SHMEM in a Lagrangian particle cluster algorithm
by: Frey, Matthias, et al.
Published: (2024)
by: Frey, Matthias, et al.
Published: (2024)
Stencil Computations on Tenstorrent Wormhole
by: Piarulli, Lorenzo, et al.
Published: (2026)
by: Piarulli, Lorenzo, et al.
Published: (2026)
Some New Approaches to MPI Implementations
by: Xiong, Yuqing
Published: (2024)
by: Xiong, Yuqing
Published: (2024)
Designing and Prototyping Extensions to MPI in MPICH
by: Zhou, Hui, et al.
Published: (2024)
by: Zhou, Hui, et al.
Published: (2024)
Concepts for designing modern C++ interfaces for MPI
by: Avans, C. Nicole, et al.
Published: (2025)
by: Avans, C. Nicole, et al.
Published: (2025)
Frustrated with MPI+Threads? Try MPIxThreads!
by: Zhou, Hui, et al.
Published: (2024)
by: Zhou, Hui, et al.
Published: (2024)
MPI-Q: A Message Communication Library for Large-Scale Classical-Quantum Heterogeneous Hybrid Distributed Computing
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
The Case for ABI Interoperability in a Fault Tolerant MPI
by: Xu, Yao, et al.
Published: (2025)
by: Xu, Yao, et al.
Published: (2025)
Parallel Spawning Strategies for Dynamic-Aware MPI Applications
by: Martín-Álvarez, Iker, et al.
Published: (2025)
by: Martín-Álvarez, Iker, et al.
Published: (2025)
Towards the Democratization and Standardization of Dynamic Resources with MPI Spawning
by: Iserte, Sergio, et al.
Published: (2026)
by: Iserte, Sergio, et al.
Published: (2026)
Do MPI Derived Datatypes Actually Help? A Single-Node Cross-Implementation Study on Shared-Memory Communication
by: Adefemi, Temitayo
Published: (2025)
by: Adefemi, Temitayo
Published: (2025)
Layout-Agnostic MPI Abstraction for Distributed Computing in Modern C++
by: Klepl, Jiří, et al.
Published: (2025)
by: Klepl, Jiří, et al.
Published: (2025)
High-Performance Parallelization of Dijkstra's Algorithm Using MPI and CUDA
by: Song, Boyang
Published: (2025)
by: Song, Boyang
Published: (2025)
On Similarity of Computational Kernels in our Codes and Proxies
by: McKinsey, Michael, et al.
Published: (2026)
by: McKinsey, Michael, et al.
Published: (2026)
Cyclic Data Streaming on GPUs for Short Range Stencils Applied to Molecular Dynamics
by: Rose, Martin, et al.
Published: (2025)
by: Rose, Martin, et al.
Published: (2025)
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
by: Iserte, Sergio, et al.
Published: (2025)
by: Iserte, Sergio, et al.
Published: (2025)
Performance of a high-order MPI-Kokkos accelerated fluid solver
by: Sporykhin, Filipp, et al.
Published: (2025)
by: Sporykhin, Filipp, et al.
Published: (2025)
MPI Malleability Validation under Replayed Real-World HPC Conditions
by: Iserte, S., et al.
Published: (2026)
by: Iserte, S., et al.
Published: (2026)
Similar Items
-
A More Scalable Sparse Dynamic Data Exchange
by: Geyko, Andrew, et al.
Published: (2023) -
To Repair or Not to Repair: Assessing Fault Resilience in MPI Stencil Applications
by: Rocco, Roberto, et al.
Published: (2024) -
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
by: Adams, Michael, et al.
Published: (2025) -
Analyzing Persistent Alltoallv RMA Implementations for High-Performance MPI Communication
by: Namugwanya, Evelyn
Published: (2026) -
Stencil Matrixization
by: Zhao, Wenxuan, et al.
Published: (2023)