Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914277802115072 |
|---|---|
| author | Kinkead, Shannon Wesley, Jackson Schonbein, Whit DeBonis, David Dosanjh, Matthew G. F. Bienz, Amanda |
| author_facet | Kinkead, Shannon Wesley, Jackson Schonbein, Whit DeBonis, David Dosanjh, Matthew G. F. Bienz, Amanda |
| contents | Performant all-to-all collective operations in MPI are critical to fast Fourier transforms, transposition, and machine learning applications. There are many existing implementations for all-to-all exchanges on emerging systems, with the achieved performance dependent on many factors, including message size, process count, architecture, and parallel system partition. This paper presents novel all-to-all algorithms for emerging many-core systems. Further, the paper presents a performance analysis against existing algorithms and system MPI, with novel algorithms achieving up to 3x speedup over system MPI at 32 nodes of state-of-the-art Sapphire Rapids systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_17606 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Scaling All-to-all Operations Across Emerging Many-Core Supercomputers Kinkead, Shannon Wesley, Jackson Schonbein, Whit DeBonis, David Dosanjh, Matthew G. F. Bienz, Amanda Distributed, Parallel, and Cluster Computing Performant all-to-all collective operations in MPI are critical to fast Fourier transforms, transposition, and machine learning applications. There are many existing implementations for all-to-all exchanges on emerging systems, with the achieved performance dependent on many factors, including message size, process count, architecture, and parallel system partition. This paper presents novel all-to-all algorithms for emerging many-core systems. Further, the paper presents a performance analysis against existing algorithms and system MPI, with novel algorithms achieving up to 3x speedup over system MPI at 32 nodes of state-of-the-art Sapphire Rapids systems. |
| title | Scaling All-to-all Operations Across Emerging Many-Core Supercomputers |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2601.17606 |