Persistent and Partitioned MPI for Stencil Communication

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Collom, Gerald, Burmark, Jason, Pearce, Olga, Bienz, Amanda
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908494279475200
author Collom, Gerald
Burmark, Jason
Pearce, Olga
Bienz, Amanda
author_facet Collom, Gerald
Burmark, Jason
Pearce, Olga
Bienz, Amanda
contents Many parallel applications rely on iterative stencil operations, whose performance are dominated by communication costs at large scales. Several MPI optimizations, such as persistent and partitioned communication, reduce overheads and improve communication efficiency through amortized setup costs and reduced synchronization of threaded sends. This paper presents the performance of stencil communication in the Comb benchmarking suite when using non blocking, persistent, and partitioned communication routines. The impact of each optimization is analyzed at various scales. Further, the paper presents an analysis of the impact of process count, thread count, and message size on partitioned communication routines. Measured timings show that persistent MPI communication can provide a speedup of up to 37% over the baseline MPI communication, and partitioned MPI communication can provide a speedup of up to 68%.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13370
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Persistent and Partitioned MPI for Stencil Communication
Collom, Gerald
Burmark, Jason
Pearce, Olga
Bienz, Amanda
Distributed, Parallel, and Cluster Computing
Many parallel applications rely on iterative stencil operations, whose performance are dominated by communication costs at large scales. Several MPI optimizations, such as persistent and partitioned communication, reduce overheads and improve communication efficiency through amortized setup costs and reduced synchronization of threaded sends. This paper presents the performance of stencil communication in the Comb benchmarking suite when using non blocking, persistent, and partitioned communication routines. The impact of each optimization is analyzed at various scales. Further, the paper presents an analysis of the impact of process count, thread count, and message size on partitioned communication routines. Measured timings show that persistent MPI communication can provide a speedup of up to 37% over the baseline MPI communication, and partitioned MPI communication can provide a speedup of up to 68%.
title Persistent and Partitioned MPI for Stencil Communication
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2508.13370