Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xue, Chunyu, Cui, Weihao, Chen, Quan, Chen, Chen, Zhao, Han, Zhang, Shulai, Wang, Linmei, Li, Yan, Xiao, Limin, Zhang, Weifeng, Yang, Jing, He, Bingsheng, Guo, Minyi
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911539163824128
author Xue, Chunyu
Cui, Weihao
Chen, Quan
Chen, Chen
Zhao, Han
Zhang, Shulai
Wang, Linmei
Li, Yan
Xiao, Limin
Zhang, Weifeng
Yang, Jing
He, Bingsheng
Guo, Minyi
author_facet Xue, Chunyu
Cui, Weihao
Chen, Quan
Chen, Chen
Zhao, Han
Zhang, Shulai
Wang, Linmei
Li, Yan
Xiao, Limin
Zhang, Weifeng
Yang, Jing
He, Bingsheng
Guo, Minyi
contents Efficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing dynamic schedulers struggle with large-model scheduling due to the mismatch between static parallelism (SP)-aware scheduling and AP-based execution, leading to cluster inefficiencies such as degraded throughput and prolonged job queuing. This paper presents Arena, a large-model training system that co-designs dynamic scheduling and adaptive parallelism to achieve high cluster efficiency. To reduce scheduling costs while improving decision quality, Arena designs low-cost, disaggregated profiling and AP-tailored, load-aware performance estimation, while unifying them by sharding the joint scheduling-parallelism optimization space via a grid abstraction. Building on this, Arena dynamically schedules profiled jobs in elasticity and heterogeneity dimensions, and executes them using efficient AP with pruned search space. Evaluated on heterogeneous testbeds and production workloads, Arena reduces job completion time by up to $49.3\%$ and improves cluster throughput by up to $1.60\times$.
format Preprint
id arxiv_https___arxiv_org_abs_2403_16125
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
Xue, Chunyu
Cui, Weihao
Chen, Quan
Chen, Chen
Zhao, Han
Zhang, Shulai
Wang, Linmei
Li, Yan
Xiao, Limin
Zhang, Weifeng
Yang, Jing
He, Bingsheng
Guo, Minyi
Distributed, Parallel, and Cluster Computing
Machine Learning
Efficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing dynamic schedulers struggle with large-model scheduling due to the mismatch between static parallelism (SP)-aware scheduling and AP-based execution, leading to cluster inefficiencies such as degraded throughput and prolonged job queuing. This paper presents Arena, a large-model training system that co-designs dynamic scheduling and adaptive parallelism to achieve high cluster efficiency. To reduce scheduling costs while improving decision quality, Arena designs low-cost, disaggregated profiling and AP-tailored, load-aware performance estimation, while unifying them by sharding the joint scheduling-parallelism optimization space via a grid abstraction. Building on this, Arena dynamically schedules profiled jobs in elasticity and heterogeneity dimensions, and executes them using efficient AP with pruned search space. Evaluated on heterogeneous testbeds and production workloads, Arena reduces job completion time by up to $49.3\%$ and improves cluster throughput by up to $1.60\times$.
title Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2403.16125