UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908368190308352 |
|---|---|
| author | Liao, Heng Liu, Bingyang Chen, Xianping Guo, Zhigang Cheng, Chuanning Wang, Jianbing Chen, Xiangyu Dong, Peng Meng, Rui Liu, Wenjie Zhou, Zhe Zhang, Ziyang Gai, Yuhang Qian, Cunle Xiong, Yi Cheng, Zhongwu Xia, Jing Ma, Yuli Chen, Xi Du, Wenhua Xiao, Shizhong Li, Chungang Qin, Yong Xiong, Liudong Yu, Zhou Chen, Lv Chen, Lei Wang, Buyun Wu, Pei Gao, Junen Li, Xiaochu He, Jian Yan, Shizhuan McColl, Bill |
| author_facet | Liao, Heng Liu, Bingyang Chen, Xianping Guo, Zhigang Cheng, Chuanning Wang, Jianbing Chen, Xiangyu Dong, Peng Meng, Rui Liu, Wenjie Zhou, Zhe Zhang, Ziyang Gai, Yuhang Qian, Cunle Xiong, Yi Cheng, Zhongwu Xia, Jing Ma, Yuli Chen, Xi Du, Wenhua Xiao, Shizhong Li, Chungang Qin, Yong Xiong, Liudong Yu, Zhou Chen, Lv Chen, Lei Wang, Buyun Wu, Pei Gao, Junen Li, Xiaochu He, Jian Yan, Shizhuan McColl, Bill |
| contents | As the Large-scale Language Models (LLMs) continue to scale, the requisite computational power and bandwidth escalate. To address this, we introduce UB-Mesh, a novel AI datacenter network architecture designed to enhance scalability, performance, cost-efficiency and availability. Unlike traditional datacenters that provide symmetrical node-to-node bandwidth, UB-Mesh employs a hierarchically localized nD-FullMesh network topology. This design fully leverages the data locality of LLM training, prioritizing short-range, direct interconnects to minimize data movement distance and reduce switch usage.
Although UB-Mesh's nD-FullMesh topology offers several theoretical advantages, its concrete architecture design, physical implementation and networking system optimization present new challenges. For the actual construction of UB-Mesh, we first design the UB-Mesh-Pod architecture, which is based on a 4D-FullMesh topology. UB-Mesh-Pod is implemented via a suite of hardware components that serve as the foundational building blocks, including specifically-designed NPU, CPU, Low-Radix-Switch (LRS), High-Radix-Switch (HRS), NICs and others. These components are interconnected via a novel Unified Bus (UB) technique, which enables flexible IO bandwidth allocation and hardware resource pooling. For networking system optimization, we propose advanced routing mechanism named All-Path-Routing (APR) to efficiently manage data traffic. These optimizations, combined with topology-aware performance enhancements and robust reliability measures like 64+1 backup design, result in 2.04x higher cost-efficiency, 7.2% higher network availability compared to traditional Clos architecture and 95%+ linearity in various LLM training tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_20377 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture Liao, Heng Liu, Bingyang Chen, Xianping Guo, Zhigang Cheng, Chuanning Wang, Jianbing Chen, Xiangyu Dong, Peng Meng, Rui Liu, Wenjie Zhou, Zhe Zhang, Ziyang Gai, Yuhang Qian, Cunle Xiong, Yi Cheng, Zhongwu Xia, Jing Ma, Yuli Chen, Xi Du, Wenhua Xiao, Shizhong Li, Chungang Qin, Yong Xiong, Liudong Yu, Zhou Chen, Lv Chen, Lei Wang, Buyun Wu, Pei Gao, Junen Li, Xiaochu He, Jian Yan, Shizhuan McColl, Bill Hardware Architecture Networking and Internet Architecture As the Large-scale Language Models (LLMs) continue to scale, the requisite computational power and bandwidth escalate. To address this, we introduce UB-Mesh, a novel AI datacenter network architecture designed to enhance scalability, performance, cost-efficiency and availability. Unlike traditional datacenters that provide symmetrical node-to-node bandwidth, UB-Mesh employs a hierarchically localized nD-FullMesh network topology. This design fully leverages the data locality of LLM training, prioritizing short-range, direct interconnects to minimize data movement distance and reduce switch usage. Although UB-Mesh's nD-FullMesh topology offers several theoretical advantages, its concrete architecture design, physical implementation and networking system optimization present new challenges. For the actual construction of UB-Mesh, we first design the UB-Mesh-Pod architecture, which is based on a 4D-FullMesh topology. UB-Mesh-Pod is implemented via a suite of hardware components that serve as the foundational building blocks, including specifically-designed NPU, CPU, Low-Radix-Switch (LRS), High-Radix-Switch (HRS), NICs and others. These components are interconnected via a novel Unified Bus (UB) technique, which enables flexible IO bandwidth allocation and hardware resource pooling. For networking system optimization, we propose advanced routing mechanism named All-Path-Routing (APR) to efficiently manage data traffic. These optimizations, combined with topology-aware performance enhancements and robust reliability measures like 64+1 backup design, result in 2.04x higher cost-efficiency, 7.2% higher network availability compared to traditional Clos architecture and 95%+ linearity in various LLM training tasks. |
| title | UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture |
| topic | Hardware Architecture Networking and Internet Architecture |
| url | https://arxiv.org/abs/2503.20377 |