Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Khalilov, Mikhail, Di Girolamo, Salvatore, Chrapek, Marcin, Nudelman, Rami, Bloch, Gil, Hoefler, Torsten
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915013780832256
author Khalilov, Mikhail
Di Girolamo, Salvatore
Chrapek, Marcin
Nudelman, Rami
Bloch, Gil
Hoefler, Torsten
author_facet Khalilov, Mikhail
Di Girolamo, Salvatore
Chrapek, Marcin
Nudelman, Rami
Bloch, Gil
Hoefler, Torsten
contents In the Fully Sharded Data Parallel (FSDP) training pipeline, collective operations can be interleaved to maximize the communication/computation overlap. In this scenario, outstanding operations such as Allgather and Reduce-Scatter can compete for the injection bandwidth and create pipeline bubbles. To address this problem, we propose a novel bandwidth-optimal Allgather collective algorithm that leverages hardware multicast. We use multicast to build a constant-time reliable Broadcast protocol, a building block for constructing an optimal Allgather schedule. Our Allgather algorithm achieves 2x traffic reduction on a 188-node testbed. To free the host side from running the protocol, we employ SmartNIC offloading. We extract the parallelism in our Allgather algorithm and map it to a SmartNIC specialized for hiding the cost of data movement. We show that our SmartNIC-offloaded collective progress engine can scale to the next generation of 1.6 Tbit/s links.
format Preprint
id arxiv_https___arxiv_org_abs_2408_13356
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
Khalilov, Mikhail
Di Girolamo, Salvatore
Chrapek, Marcin
Nudelman, Rami
Bloch, Gil
Hoefler, Torsten
Distributed, Parallel, and Cluster Computing
In the Fully Sharded Data Parallel (FSDP) training pipeline, collective operations can be interleaved to maximize the communication/computation overlap. In this scenario, outstanding operations such as Allgather and Reduce-Scatter can compete for the injection bandwidth and create pipeline bubbles. To address this problem, we propose a novel bandwidth-optimal Allgather collective algorithm that leverages hardware multicast. We use multicast to build a constant-time reliable Broadcast protocol, a building block for constructing an optimal Allgather schedule. Our Allgather algorithm achieves 2x traffic reduction on a 188-node testbed. To free the host side from running the protocol, we employ SmartNIC offloading. We extract the parallelism in our Allgather algorithm and map it to a SmartNIC specialized for hiding the cost of data movement. We show that our SmartNIC-offloaded collective progress engine can scale to the next generation of 1.6 Tbit/s links.
title Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2408.13356