GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Jiaang, Xu, Shenglin, Qian, Shiyou, Yang, Dingyu, Wang, Kangjin, Liao, Chenzhi, Yu, Yinghao, Hua, Qin, Hu, Hanwen, Wang, Qi, Wu, Wenchao, Bao, Dongqing, Lu, Tianyu, Cao, Jian, Xue, Guangtao, Yang, Guodong, Zhang, Liping, Chen, Gang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908538472759296
author Duan, Jiaang
Xu, Shenglin
Qian, Shiyou
Yang, Dingyu
Wang, Kangjin
Liao, Chenzhi
Yu, Yinghao
Hua, Qin
Hu, Hanwen
Wang, Qi
Wu, Wenchao
Bao, Dongqing
Lu, Tianyu
Cao, Jian
Xue, Guangtao
Yang, Guodong
Zhang, Liping
Chen, Gang
author_facet Duan, Jiaang
Xu, Shenglin
Qian, Shiyou
Yang, Dingyu
Wang, Kangjin
Liao, Chenzhi
Yu, Yinghao
Hua, Qin
Hu, Hanwen
Wang, Qi
Wu, Wenchao
Bao, Dongqing
Lu, Tianyu
Cao, Jian
Xue, Guangtao
Yang, Guodong
Zhang, Liping
Chen, Gang
contents The surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cloud providers employ spot instances to reduce costs for low-priority (LP) tasks, existing schedulers still grapple with high eviction rates and lengthy queuing times. To address these limitations, we present GFS, a novel preemptive scheduling framework that enhances service-level objective (SLO) compliance for high-priority (HP) tasks while minimizing preemptions to LP tasks. Firstly, GFS utilizes a lightweight forecasting model that predicts GPU demand among different tenants, enabling proactive resource management. Secondly, GFS employs a dynamic allocation mechanism to adjust the spot quota for LP tasks with guaranteed durations. Lastly, GFS incorporates a preemptive scheduling policy that prioritizes HP tasks while minimizing the impact on LP tasks. We demonstrate the effectiveness of GFS through both real-world implementation and simulations. The results show that GFS reduces eviction rates by 33.0\%, and cuts queuing delays by 44.1\% for LP tasks. Furthermore, GFS enhances the GPU allocation rate by up to 22.8\% in real production clusters. In a production cluster of more than 10,000 GPUs, GFS yields roughly \$459,715 in monthly benefits.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11134
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
Duan, Jiaang
Xu, Shenglin
Qian, Shiyou
Yang, Dingyu
Wang, Kangjin
Liao, Chenzhi
Yu, Yinghao
Hua, Qin
Hu, Hanwen
Wang, Qi
Wu, Wenchao
Bao, Dongqing
Lu, Tianyu
Cao, Jian
Xue, Guangtao
Yang, Guodong
Zhang, Liping
Chen, Gang
Distributed, Parallel, and Cluster Computing
The surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cloud providers employ spot instances to reduce costs for low-priority (LP) tasks, existing schedulers still grapple with high eviction rates and lengthy queuing times. To address these limitations, we present GFS, a novel preemptive scheduling framework that enhances service-level objective (SLO) compliance for high-priority (HP) tasks while minimizing preemptions to LP tasks. Firstly, GFS utilizes a lightweight forecasting model that predicts GPU demand among different tenants, enabling proactive resource management. Secondly, GFS employs a dynamic allocation mechanism to adjust the spot quota for LP tasks with guaranteed durations. Lastly, GFS incorporates a preemptive scheduling policy that prioritizes HP tasks while minimizing the impact on LP tasks. We demonstrate the effectiveness of GFS through both real-world implementation and simulations. The results show that GFS reduces eviction rates by 33.0\%, and cuts queuing delays by 44.1\% for LP tasks. Furthermore, GFS enhances the GPU allocation rate by up to 22.8\% in real production clusters. In a production cluster of more than 10,000 GPUs, GFS yields roughly \$459,715 in monthly benefits.
title GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.11134