ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Minghao, Golden, Alicia, Hsia, Samuel, Kuchnik, Michael, Gangidi, Adi, Zhang, Xu, Shetty, Ashmitha Jeevaraj, DeVito, Zachary, Chu, Weiwei, He, Dong, Zhang, Haoci, Hao, Yuchen, Pang, Ruoming, Zeng, James Hongyi, Zhang, Ying, Yu, Minlan, Wu, Carole-Jean
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916041413623808
author Li, Minghao
Golden, Alicia
Hsia, Samuel
Kuchnik, Michael
Gangidi, Adi
Zhang, Xu
Shetty, Ashmitha Jeevaraj
DeVito, Zachary
Chu, Weiwei
He, Dong
Zhang, Haoci
Hao, Yuchen
Pang, Ruoming
Zeng, James Hongyi
Zhang, Ying
Yu, Minlan
Wu, Carole-Jean
author_facet Li, Minghao
Golden, Alicia
Hsia, Samuel
Kuchnik, Michael
Gangidi, Adi
Zhang, Xu
Shetty, Ashmitha Jeevaraj
DeVito, Zachary
Chu, Weiwei
He, Dong
Zhang, Haoci
Hao, Yuchen
Pang, Ruoming
Zeng, James Hongyi
Zhang, Ying
Yu, Minlan
Wu, Carole-Jean
contents The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to such paradigm as "scale-across" training. As infrastructure expands, the system design space becomes increasingly intricate, encompassing new model architectures, hardware heterogeneity, and evolving communication patterns. Drawing from Meta's production experience, we highlight the complexities of deploying training jobs across a few data centers housing hundreds of thousands of GPUs. To accelerate exploration of the large design space and to enable efficient training for frontier model development, we conduct in-depth characterization of three key design dimensions: parallelism placement, parallelism scheduling, and network layer technologies. We then propose ScaleAcross Explorer, an optimizer that considers the interplay of design dimensions and holistically optimizes scale-across training. Testbed experiments and simulations demonstrate up to 64.62% training speedups over production configuration and up to 37.59% training speedups over the state-of-the-art baseline across a wide range of design points.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24326
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
Li, Minghao
Golden, Alicia
Hsia, Samuel
Kuchnik, Michael
Gangidi, Adi
Zhang, Xu
Shetty, Ashmitha Jeevaraj
DeVito, Zachary
Chu, Weiwei
He, Dong
Zhang, Haoci
Hao, Yuchen
Pang, Ruoming
Zeng, James Hongyi
Zhang, Ying
Yu, Minlan
Wu, Carole-Jean
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Networking and Internet Architecture
The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to such paradigm as "scale-across" training. As infrastructure expands, the system design space becomes increasingly intricate, encompassing new model architectures, hardware heterogeneity, and evolving communication patterns. Drawing from Meta's production experience, we highlight the complexities of deploying training jobs across a few data centers housing hundreds of thousands of GPUs. To accelerate exploration of the large design space and to enable efficient training for frontier model development, we conduct in-depth characterization of three key design dimensions: parallelism placement, parallelism scheduling, and network layer technologies. We then propose ScaleAcross Explorer, an optimizer that considers the interplay of design dimensions and holistically optimizes scale-across training. Testbed experiments and simulations demonstrate up to 64.62% training speedups over production configuration and up to 37.59% training speedups over the state-of-the-art baseline across a wide range of design points.
title ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Networking and Internet Architecture
url https://arxiv.org/abs/2605.24326