Leaf-centric Logical Topology Design for OCS-based GPU Clusters

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Han, Xinchi, Jiang, Weihao, Mao, Yingming, Liu, Yike, Liu, Zhuoran, Lv, Yongxi, Cao, Peirui, Liu, Zhuotao, Liu, Ximeng, Wang, Xinbing, Wu, Changbo, Zhu, Zihan, Dongchao, Wu, Jian, Yang, Zhanbang, Zhang, Chen, Yuansen, Zhao, Shizhen
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912989726113792
author Han, Xinchi
Jiang, Weihao
Mao, Yingming
Liu, Yike
Liu, Zhuoran
Lv, Yongxi
Cao, Peirui
Liu, Zhuotao
Liu, Ximeng
Wang, Xinbing
Wu, Changbo
Zhu, Zihan
Dongchao, Wu
Jian, Yang
Zhanbang, Zhang
Chen, Yuansen
Zhao, Shizhen
author_facet Han, Xinchi
Jiang, Weihao
Mao, Yingming
Liu, Yike
Liu, Zhuoran
Lv, Yongxi
Cao, Peirui
Liu, Zhuotao
Liu, Ximeng
Wang, Xinbing
Wu, Changbo
Zhu, Zihan
Dongchao, Wu
Jian, Yang
Zhanbang, Zhang
Chen, Yuansen
Zhao, Shizhen
contents Recent years have witnessed the growing deployment of optical circuit switches (OCS) in commercial GPU clusters (e.g., Google A3 GPU cluster) optimized for machine learning (ML) workloads. Such clusters adopt a three-tier leaf-spine-OCS topology, servers attach to leaf-layer electronic packet switches (EPSes); these leaf switches aggregate into spine-layer EPSes to form a Pod; and multiple Pods are interconnected via core-layer OCSes. Unlike EPSes, OCSes only support circuit-based paths between directly connected spine switches, potentially inducing a phenomenon termed routing polarization, which refers to the scenario where the bandwidth requirements between specific pairs of Pods are unevenly fulfilled through links among different spine switches. The resulting imbalance induces traffic contention and bottlenecks on specific leaf-to-spine links, ultimately reducing ML training throughput. To mitigate this issue, we introduce a leaf-centric paradigm to ensure traffic originating from the same leaf switch is evenly distributed across multiple spine switches with balanced loads. Through rigorous theoretical analysis, we establish a sufficient condition for avoiding routing polarization and propose a corresponding logical topology design algorithm with polynomial-time complexity. Large-scale simulations validate up to 19.27% throughput improvement and a 99.16% reduction in logical topology computation overhead compared to Mixed Integer Programming (MIP)-based methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_28168
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Leaf-centric Logical Topology Design for OCS-based GPU Clusters
Han, Xinchi
Jiang, Weihao
Mao, Yingming
Liu, Yike
Liu, Zhuoran
Lv, Yongxi
Cao, Peirui
Liu, Zhuotao
Liu, Ximeng
Wang, Xinbing
Wu, Changbo
Zhu, Zihan
Dongchao, Wu
Jian, Yang
Zhanbang, Zhang
Chen, Yuansen
Zhao, Shizhen
Networking and Internet Architecture
Recent years have witnessed the growing deployment of optical circuit switches (OCS) in commercial GPU clusters (e.g., Google A3 GPU cluster) optimized for machine learning (ML) workloads. Such clusters adopt a three-tier leaf-spine-OCS topology, servers attach to leaf-layer electronic packet switches (EPSes); these leaf switches aggregate into spine-layer EPSes to form a Pod; and multiple Pods are interconnected via core-layer OCSes. Unlike EPSes, OCSes only support circuit-based paths between directly connected spine switches, potentially inducing a phenomenon termed routing polarization, which refers to the scenario where the bandwidth requirements between specific pairs of Pods are unevenly fulfilled through links among different spine switches. The resulting imbalance induces traffic contention and bottlenecks on specific leaf-to-spine links, ultimately reducing ML training throughput. To mitigate this issue, we introduce a leaf-centric paradigm to ensure traffic originating from the same leaf switch is evenly distributed across multiple spine switches with balanced loads. Through rigorous theoretical analysis, we establish a sufficient condition for avoiding routing polarization and propose a corresponding logical topology design algorithm with polynomial-time complexity. Large-scale simulations validate up to 19.27% throughput improvement and a 99.16% reduction in logical topology computation overhead compared to Mixed Integer Programming (MIP)-based methods.
title Leaf-centric Logical Topology Design for OCS-based GPU Clusters
topic Networking and Internet Architecture
url https://arxiv.org/abs/2603.28168