Scaling All-to-all Operations Across Emerging Many-Core Supercomputers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kinkead, Shannon, Wesley, Jackson, Schonbein, Whit, DeBonis, David, Dosanjh, Matthew G. F., Bienz, Amanda
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914277802115072
author Kinkead, Shannon
Wesley, Jackson
Schonbein, Whit
DeBonis, David
Dosanjh, Matthew G. F.
Bienz, Amanda
author_facet Kinkead, Shannon
Wesley, Jackson
Schonbein, Whit
DeBonis, David
Dosanjh, Matthew G. F.
Bienz, Amanda
contents Performant all-to-all collective operations in MPI are critical to fast Fourier transforms, transposition, and machine learning applications. There are many existing implementations for all-to-all exchanges on emerging systems, with the achieved performance dependent on many factors, including message size, process count, architecture, and parallel system partition. This paper presents novel all-to-all algorithms for emerging many-core systems. Further, the paper presents a performance analysis against existing algorithms and system MPI, with novel algorithms achieving up to 3x speedup over system MPI at 32 nodes of state-of-the-art Sapphire Rapids systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17606
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
Kinkead, Shannon
Wesley, Jackson
Schonbein, Whit
DeBonis, David
Dosanjh, Matthew G. F.
Bienz, Amanda
Distributed, Parallel, and Cluster Computing
Performant all-to-all collective operations in MPI are critical to fast Fourier transforms, transposition, and machine learning applications. There are many existing implementations for all-to-all exchanges on emerging systems, with the achieved performance dependent on many factors, including message size, process count, architecture, and parallel system partition. This paper presents novel all-to-all algorithms for emerging many-core systems. Further, the paper presents a performance analysis against existing algorithms and system MPI, with novel algorithms achieving up to 3x speedup over system MPI at 32 nodes of state-of-the-art Sapphire Rapids systems.
title Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2601.17606