PCCL: Photonic circuit-switched collective communication for distributed ML

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Abhishek Vijaya, Devraj, Arjun, Singh, Rachee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916957247242240
author Kumar, Abhishek Vijaya
Devraj, Arjun
Singh, Rachee
author_facet Kumar, Abhishek Vijaya
Devraj, Arjun
Singh, Rachee
contents Modern distributed ML suffers from a fundamental gap between the theoretical and realized performance of collective communication algorithms due to congestion and hop-count induced dilation in practical GPU clusters. We present PCCL, a Photonic Collective Communication Library that reconfigures the network topology to match the communication patterns of collective algorithms, thereby eliminating congestion and dilation by creating direct, contention-free circuits between communicating GPUs. Unlike prior approaches that synthesize algorithms for specific network topologies and collectives, PCCL generalizes to any collective primitive and any topology by adapting the network to match each algorithm's communication pattern. PCCL's key innovation lies in its hardware-agnostic optimization framework that intelligently decides when to reconfigure based on the trade-off between network reconfiguration delay and congestion/dilation costs, making it practical across different optical hardware with varying switching speeds. Our evaluation demonstrates that PCCL achieves up to 3X speedup over state-of-the-art algorithms on 128 GPUs across various workloads, buffer sizes, and topologies, translating to a 1.3X speedup in end-to-end training throughput.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15450
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PCCL: Photonic circuit-switched collective communication for distributed ML
Kumar, Abhishek Vijaya
Devraj, Arjun
Singh, Rachee
Distributed, Parallel, and Cluster Computing
Modern distributed ML suffers from a fundamental gap between the theoretical and realized performance of collective communication algorithms due to congestion and hop-count induced dilation in practical GPU clusters. We present PCCL, a Photonic Collective Communication Library that reconfigures the network topology to match the communication patterns of collective algorithms, thereby eliminating congestion and dilation by creating direct, contention-free circuits between communicating GPUs. Unlike prior approaches that synthesize algorithms for specific network topologies and collectives, PCCL generalizes to any collective primitive and any topology by adapting the network to match each algorithm's communication pattern. PCCL's key innovation lies in its hardware-agnostic optimization framework that intelligently decides when to reconfigure based on the trade-off between network reconfiguration delay and congestion/dilation costs, making it practical across different optical hardware with varying switching speeds. Our evaluation demonstrates that PCCL achieves up to 3X speedup over state-of-the-art algorithms on 128 GPUs across various workloads, buffer sizes, and topologies, translating to a 1.3X speedup in end-to-end training throughput.
title PCCL: Photonic circuit-switched collective communication for distributed ML
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.15450