Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911539163824128 |
|---|---|
| author | Xue, Chunyu Cui, Weihao Chen, Quan Chen, Chen Zhao, Han Zhang, Shulai Wang, Linmei Li, Yan Xiao, Limin Zhang, Weifeng Yang, Jing He, Bingsheng Guo, Minyi |
| author_facet | Xue, Chunyu Cui, Weihao Chen, Quan Chen, Chen Zhao, Han Zhang, Shulai Wang, Linmei Li, Yan Xiao, Limin Zhang, Weifeng Yang, Jing He, Bingsheng Guo, Minyi |
| contents | Efficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing dynamic schedulers struggle with large-model scheduling due to the mismatch between static parallelism (SP)-aware scheduling and AP-based execution, leading to cluster inefficiencies such as degraded throughput and prolonged job queuing. This paper presents Arena, a large-model training system that co-designs dynamic scheduling and adaptive parallelism to achieve high cluster efficiency. To reduce scheduling costs while improving decision quality, Arena designs low-cost, disaggregated profiling and AP-tailored, load-aware performance estimation, while unifying them by sharding the joint scheduling-parallelism optimization space via a grid abstraction. Building on this, Arena dynamically schedules profiled jobs in elasticity and heterogeneity dimensions, and executes them using efficient AP with pruned search space. Evaluated on heterogeneous testbeds and production workloads, Arena reduces job completion time by up to $49.3\%$ and improves cluster throughput by up to $1.60\times$. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_16125 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design Xue, Chunyu Cui, Weihao Chen, Quan Chen, Chen Zhao, Han Zhang, Shulai Wang, Linmei Li, Yan Xiao, Limin Zhang, Weifeng Yang, Jing He, Bingsheng Guo, Minyi Distributed, Parallel, and Cluster Computing Machine Learning Efficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing dynamic schedulers struggle with large-model scheduling due to the mismatch between static parallelism (SP)-aware scheduling and AP-based execution, leading to cluster inefficiencies such as degraded throughput and prolonged job queuing. This paper presents Arena, a large-model training system that co-designs dynamic scheduling and adaptive parallelism to achieve high cluster efficiency. To reduce scheduling costs while improving decision quality, Arena designs low-cost, disaggregated profiling and AP-tailored, load-aware performance estimation, while unifying them by sharding the joint scheduling-parallelism optimization space via a grid abstraction. Building on this, Arena dynamically schedules profiled jobs in elasticity and heterogeneity dimensions, and executes them using efficient AP with pruned search space. Evaluated on heterogeneous testbeds and production workloads, Arena reduces job completion time by up to $49.3\%$ and improves cluster throughput by up to $1.60\times$. |
| title | Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design |
| topic | Distributed, Parallel, and Cluster Computing Machine Learning |
| url | https://arxiv.org/abs/2403.16125 |