DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Jinquan, Liao, Xiaojian, Liu, Xuzhao, Suo, Jiashun, Huo, Zhisheng, Zhang, Chenhao, Xu, Xiangrong, Shen, Runnan, Xie, Xilong, Xiao, Limin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909624515428352
author Wang, Jinquan
Liao, Xiaojian
Liu, Xuzhao
Suo, Jiashun
Huo, Zhisheng
Zhang, Chenhao
Xu, Xiangrong
Shen, Runnan
Xie, Xilong
Xiao, Limin
author_facet Wang, Jinquan
Liao, Xiaojian
Liu, Xuzhao
Suo, Jiashun
Huo, Zhisheng
Zhang, Chenhao
Xu, Xiangrong
Shen, Runnan
Xie, Xilong
Xiao, Limin
contents Most existing training systems focus on a single region. In contrast, we envision that cross-region training offers more flexible GPU resource allocation and yields significant potential. However, the hierarchical cluster topology and unstable networks in the cloud-edge-end (CEE) environment, a typical cross-region scenario, pose substantial challenges to building an efficient and autonomous model training system. We propose DeepCEE, a geo-distributed model training system tailored for heterogeneous GPUs and networks in CEE environments. DeepCEE adopts a communication-centric design philosophy to tackle challenges arising from slow and unstable inter-region networks. It begins with a heterogeneous device profiler that identifies and groups devices based on both network and compute characteristics. Leveraging device groups, DeepCEE implements compact, zero-bubble pipeline parallelism, automatically deriving optimal parallel strategies. To further adapt to runtime variability, DeepCEE integrates a dynamic environment adapter that reacts to network fluctuations. Extensive evaluations demonstrate that DeepCEE achieves 1.3-2.8x higher training throughput compared to widely used and SOTA training systems.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15536
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks
Wang, Jinquan
Liao, Xiaojian
Liu, Xuzhao
Suo, Jiashun
Huo, Zhisheng
Zhang, Chenhao
Xu, Xiangrong
Shen, Runnan
Xie, Xilong
Xiao, Limin
Systems and Control
Distributed, Parallel, and Cluster Computing
Most existing training systems focus on a single region. In contrast, we envision that cross-region training offers more flexible GPU resource allocation and yields significant potential. However, the hierarchical cluster topology and unstable networks in the cloud-edge-end (CEE) environment, a typical cross-region scenario, pose substantial challenges to building an efficient and autonomous model training system. We propose DeepCEE, a geo-distributed model training system tailored for heterogeneous GPUs and networks in CEE environments. DeepCEE adopts a communication-centric design philosophy to tackle challenges arising from slow and unstable inter-region networks. It begins with a heterogeneous device profiler that identifies and groups devices based on both network and compute characteristics. Leveraging device groups, DeepCEE implements compact, zero-bubble pipeline parallelism, automatically deriving optimal parallel strategies. To further adapt to runtime variability, DeepCEE integrates a dynamic environment adapter that reacts to network fluctuations. Extensive evaluations demonstrate that DeepCEE achieves 1.3-2.8x higher training throughput compared to widely used and SOTA training systems.
title DeepCEE: Efficient Cross-Region Model Distributed Training System under Heterogeneous GPUs and Networks
topic Systems and Control
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2505.15536