Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Basu, Prithwish, Zhao, Liangyu, Fantl, Jason, Pal, Siddharth, Krishnamurthy, Arvind, Khoury, Joud
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911853514326016
author Basu, Prithwish
Zhao, Liangyu
Fantl, Jason
Pal, Siddharth
Krishnamurthy, Arvind
Khoury, Joud
author_facet Basu, Prithwish
Zhao, Liangyu
Fantl, Jason
Pal, Siddharth
Krishnamurthy, Arvind
Khoury, Joud
contents The all-to-all collective communications primitive is widely used in machine learning (ML) and high performance computing (HPC) workloads, and optimizing its performance is of interest to both ML and HPC communities. All-to-all is a particularly challenging workload that can severely strain the underlying interconnect bandwidth at scale. This paper takes a holistic approach to optimize the performance of all-to-all collective communications on supercomputer-scale direct-connect interconnects. We address several algorithmic and practical challenges in developing efficient and bandwidth-optimal all-to-all schedules for any topology and lowering the schedules to various runtimes and interconnect technologies. We also propose a novel topology that delivers near-optimal all-to-all performance.
format Preprint
id arxiv_https___arxiv_org_abs_2309_13541
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
Basu, Prithwish
Zhao, Liangyu
Fantl, Jason
Pal, Siddharth
Krishnamurthy, Arvind
Khoury, Joud
Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
The all-to-all collective communications primitive is widely used in machine learning (ML) and high performance computing (HPC) workloads, and optimizing its performance is of interest to both ML and HPC communities. All-to-all is a particularly challenging workload that can severely strain the underlying interconnect bandwidth at scale. This paper takes a holistic approach to optimize the performance of all-to-all collective communications on supercomputer-scale direct-connect interconnects. We address several algorithmic and practical challenges in developing efficient and bandwidth-optimal all-to-all schedules for any topology and lowering the schedules to various runtimes and interconnect technologies. We also propose a novel topology that delivers near-optimal all-to-all performance.
title Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
topic Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
url https://arxiv.org/abs/2309.13541