UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liao, Heng, Liu, Bingyang, Chen, Xianping, Guo, Zhigang, Cheng, Chuanning, Wang, Jianbing, Chen, Xiangyu, Dong, Peng, Meng, Rui, Liu, Wenjie, Zhou, Zhe, Zhang, Ziyang, Gai, Yuhang, Qian, Cunle, Xiong, Yi, Cheng, Zhongwu, Xia, Jing, Ma, Yuli, Chen, Xi, Du, Wenhua, Xiao, Shizhong, Li, Chungang, Qin, Yong, Xiong, Liudong, Yu, Zhou, Chen, Lv, Chen, Lei, Wang, Buyun, Wu, Pei, Gao, Junen, Li, Xiaochu, He, Jian, Yan, Shizhuan, McColl, Bill
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908368190308352
author Liao, Heng
Liu, Bingyang
Chen, Xianping
Guo, Zhigang
Cheng, Chuanning
Wang, Jianbing
Chen, Xiangyu
Dong, Peng
Meng, Rui
Liu, Wenjie
Zhou, Zhe
Zhang, Ziyang
Gai, Yuhang
Qian, Cunle
Xiong, Yi
Cheng, Zhongwu
Xia, Jing
Ma, Yuli
Chen, Xi
Du, Wenhua
Xiao, Shizhong
Li, Chungang
Qin, Yong
Xiong, Liudong
Yu, Zhou
Chen, Lv
Chen, Lei
Wang, Buyun
Wu, Pei
Gao, Junen
Li, Xiaochu
He, Jian
Yan, Shizhuan
McColl, Bill
author_facet Liao, Heng
Liu, Bingyang
Chen, Xianping
Guo, Zhigang
Cheng, Chuanning
Wang, Jianbing
Chen, Xiangyu
Dong, Peng
Meng, Rui
Liu, Wenjie
Zhou, Zhe
Zhang, Ziyang
Gai, Yuhang
Qian, Cunle
Xiong, Yi
Cheng, Zhongwu
Xia, Jing
Ma, Yuli
Chen, Xi
Du, Wenhua
Xiao, Shizhong
Li, Chungang
Qin, Yong
Xiong, Liudong
Yu, Zhou
Chen, Lv
Chen, Lei
Wang, Buyun
Wu, Pei
Gao, Junen
Li, Xiaochu
He, Jian
Yan, Shizhuan
McColl, Bill
contents As the Large-scale Language Models (LLMs) continue to scale, the requisite computational power and bandwidth escalate. To address this, we introduce UB-Mesh, a novel AI datacenter network architecture designed to enhance scalability, performance, cost-efficiency and availability. Unlike traditional datacenters that provide symmetrical node-to-node bandwidth, UB-Mesh employs a hierarchically localized nD-FullMesh network topology. This design fully leverages the data locality of LLM training, prioritizing short-range, direct interconnects to minimize data movement distance and reduce switch usage. Although UB-Mesh's nD-FullMesh topology offers several theoretical advantages, its concrete architecture design, physical implementation and networking system optimization present new challenges. For the actual construction of UB-Mesh, we first design the UB-Mesh-Pod architecture, which is based on a 4D-FullMesh topology. UB-Mesh-Pod is implemented via a suite of hardware components that serve as the foundational building blocks, including specifically-designed NPU, CPU, Low-Radix-Switch (LRS), High-Radix-Switch (HRS), NICs and others. These components are interconnected via a novel Unified Bus (UB) technique, which enables flexible IO bandwidth allocation and hardware resource pooling. For networking system optimization, we propose advanced routing mechanism named All-Path-Routing (APR) to efficiently manage data traffic. These optimizations, combined with topology-aware performance enhancements and robust reliability measures like 64+1 backup design, result in 2.04x higher cost-efficiency, 7.2% higher network availability compared to traditional Clos architecture and 95%+ linearity in various LLM training tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2503_20377
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture
Liao, Heng
Liu, Bingyang
Chen, Xianping
Guo, Zhigang
Cheng, Chuanning
Wang, Jianbing
Chen, Xiangyu
Dong, Peng
Meng, Rui
Liu, Wenjie
Zhou, Zhe
Zhang, Ziyang
Gai, Yuhang
Qian, Cunle
Xiong, Yi
Cheng, Zhongwu
Xia, Jing
Ma, Yuli
Chen, Xi
Du, Wenhua
Xiao, Shizhong
Li, Chungang
Qin, Yong
Xiong, Liudong
Yu, Zhou
Chen, Lv
Chen, Lei
Wang, Buyun
Wu, Pei
Gao, Junen
Li, Xiaochu
He, Jian
Yan, Shizhuan
McColl, Bill
Hardware Architecture
Networking and Internet Architecture
As the Large-scale Language Models (LLMs) continue to scale, the requisite computational power and bandwidth escalate. To address this, we introduce UB-Mesh, a novel AI datacenter network architecture designed to enhance scalability, performance, cost-efficiency and availability. Unlike traditional datacenters that provide symmetrical node-to-node bandwidth, UB-Mesh employs a hierarchically localized nD-FullMesh network topology. This design fully leverages the data locality of LLM training, prioritizing short-range, direct interconnects to minimize data movement distance and reduce switch usage. Although UB-Mesh's nD-FullMesh topology offers several theoretical advantages, its concrete architecture design, physical implementation and networking system optimization present new challenges. For the actual construction of UB-Mesh, we first design the UB-Mesh-Pod architecture, which is based on a 4D-FullMesh topology. UB-Mesh-Pod is implemented via a suite of hardware components that serve as the foundational building blocks, including specifically-designed NPU, CPU, Low-Radix-Switch (LRS), High-Radix-Switch (HRS), NICs and others. These components are interconnected via a novel Unified Bus (UB) technique, which enables flexible IO bandwidth allocation and hardware resource pooling. For networking system optimization, we propose advanced routing mechanism named All-Path-Routing (APR) to efficiently manage data traffic. These optimizations, combined with topology-aware performance enhancements and robust reliability measures like 64+1 backup design, result in 2.04x higher cost-efficiency, 7.2% higher network availability compared to traditional Clos architecture and 95%+ linearity in various LLM training tasks.
title UB-Mesh: a Hierarchically Localized nD-FullMesh Datacenter Network Architecture
topic Hardware Architecture
Networking and Internet Architecture
url https://arxiv.org/abs/2503.20377