Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Yujie, Zhu, Shenhan, Fu, Fangcheng, Miao, Xupeng, Zhang, Jie, Zhu, Juan, Hong, Fan, Li, Yong, Cui, Bin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909487177138176
author Wang, Yujie
Zhu, Shenhan
Fu, Fangcheng
Miao, Xupeng
Zhang, Jie
Zhu, Juan
Hong, Fan
Li, Yong
Cui, Bin
author_facet Wang, Yujie
Zhu, Shenhan
Fu, Fangcheng
Miao, Xupeng
Zhang, Jie
Zhu, Juan
Hong, Fan
Li, Yong
Cui, Bin
contents Recent foundation models are capable of handling multiple tasks and multiple data modalities with the unified base model structure and several specialized model components. However, efficient training of such multi-task (MT) multi-modal (MM) models poses significant system challenges due to the sophisticated model architecture and the heterogeneous workloads of different tasks and modalities. In this paper, we propose Spindle, a brand new training system tailored for resource-efficient and high-performance training of MT MM models via wavefront scheduling. The key idea of Spindle is to decompose the model execution into waves and address the joint optimization problem sequentially, including both heterogeneity-aware workload parallelization and dependency-driven execution scheduling. We build our system and evaluate it on various MT MM models. Experiments demonstrate the superior performance and efficiency of Spindle, with speedup ratio up to 71% compared to state-of-the-art training systems.
format Preprint
id arxiv_https___arxiv_org_abs_2409_03365
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling
Wang, Yujie
Zhu, Shenhan
Fu, Fangcheng
Miao, Xupeng
Zhang, Jie
Zhu, Juan
Hong, Fan
Li, Yong
Cui, Bin
Distributed, Parallel, and Cluster Computing
Machine Learning
Recent foundation models are capable of handling multiple tasks and multiple data modalities with the unified base model structure and several specialized model components. However, efficient training of such multi-task (MT) multi-modal (MM) models poses significant system challenges due to the sophisticated model architecture and the heterogeneous workloads of different tasks and modalities. In this paper, we propose Spindle, a brand new training system tailored for resource-efficient and high-performance training of MT MM models via wavefront scheduling. The key idea of Spindle is to decompose the model execution into waves and address the joint optimization problem sequentially, including both heterogeneity-aware workload parallelization and dependency-driven execution scheduling. We build our system and evaluate it on various MT MM models. Experiments demonstrate the superior performance and efficiency of Spindle, with speedup ratio up to 71% compared to state-of-the-art training systems.
title Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2409.03365