Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chang, Zihan, Xiao, Sheng, He, Shuibing, Yang, Siling, Pan, Zhe, Li, Dong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929639957463040
author Chang, Zihan
Xiao, Sheng
He, Shuibing
Yang, Siling
Pan, Zhe
Li, Dong
author_facet Chang, Zihan
Xiao, Sheng
He, Shuibing
Yang, Siling
Pan, Zhe
Li, Dong
contents Existing work only effective on a given number of GPUs, often neglecting the complexities involved in manually determining the specific types and quantities of GPUs needed, which can be a significant burden for developers. To address this issue, we propose Frenzy, a memory-aware serverless computing method for heterogeneous GPU clusters. Frenzy allows users to submit models without worrying about underlying hardware resources. First, Frenzy predicts the required number and type of GPUs by estimating the GPU memory usage of the LLM. Then, it employs a low-overhead heterogeneity-aware scheduling method to optimize training efficiency. We validated Frenzy's performance by conducting multi-task LLM training tests on a heterogeneous GPU cluster with three different GPU types. The results show that Frenzy's memory usage prediction accuracy exceeds 92\%, the scheduling overhead is reduced by 10 times, and it reduces the average job completion time by 12\% to 18\% compared to state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14479
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
Chang, Zihan
Xiao, Sheng
He, Shuibing
Yang, Siling
Pan, Zhe
Li, Dong
Distributed, Parallel, and Cluster Computing
Existing work only effective on a given number of GPUs, often neglecting the complexities involved in manually determining the specific types and quantities of GPUs needed, which can be a significant burden for developers. To address this issue, we propose Frenzy, a memory-aware serverless computing method for heterogeneous GPU clusters. Frenzy allows users to submit models without worrying about underlying hardware resources. First, Frenzy predicts the required number and type of GPUs by estimating the GPU memory usage of the LLM. Then, it employs a low-overhead heterogeneity-aware scheduling method to optimize training efficiency. We validated Frenzy's performance by conducting multi-task LLM training tests on a heterogeneous GPU cluster with three different GPU types. The results show that Frenzy's memory usage prediction accuracy exceeds 92\%, the scheduling overhead is reduced by 10 times, and it reduces the average job completion time by 12\% to 18\% compared to state-of-the-art methods.
title Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2412.14479