GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Jiaang, Xu, Shenglin, Qian, Shiyou, Yang, Dingyu, Wang, Kangjin, Liao, Chenzhi, Yu, Yinghao, Hua, Qin, Hu, Hanwen, Wang, Qi, Wu, Wenchao, Bao, Dongqing, Lu, Tianyu, Cao, Jian, Xue, Guangtao, Yang, Guodong, Zhang, Liping, Chen, Gang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
by: Duan, Jiaang, et al.
Published: (2024)
by: Duan, Jiaang, et al.
Published: (2024)
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
by: Yi, Qingao, et al.
Published: (2025)
by: Yi, Qingao, et al.
Published: (2025)
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
by: Zhang, Kaixuan, et al.
Published: (2026)
by: Zhang, Kaixuan, et al.
Published: (2026)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
by: Yang, Dingyu, et al.
Published: (2026)
by: Yang, Dingyu, et al.
Published: (2026)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
by: Hua, Qin, et al.
Published: (2024)
by: Hua, Qin, et al.
Published: (2024)
Mitigating Interference of Microservices with a Scoring Mechanism in Large-scale Clusters
by: Yang, Dingyu, et al.
Published: (2024)
by: Yang, Dingyu, et al.
Published: (2024)
Atys: An Efficient Profiling Framework for Identifying Hotspot Functions in Large-scale Cloud Microservices
by: Sun, Jiaqi, et al.
Published: (2025)
by: Sun, Jiaqi, et al.
Published: (2025)
PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance Prediction
by: Zhang, Kaixuan, et al.
Published: (2026)
by: Zhang, Kaixuan, et al.
Published: (2026)
InstGenIE: Generative Image Editing Made Efficient with Mask-aware Caching and Scheduling
by: Jiang, Xiaoxiao, et al.
Published: (2025)
by: Jiang, Xiaoxiao, et al.
Published: (2025)
Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications
by: Wang, Yuxiao, et al.
Published: (2025)
by: Wang, Yuxiao, et al.
Published: (2025)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)
by: Wu, Tianyuan, et al.
Published: (2024)
Hummingbird: SLO-Oriented GPU Preemption at Microsecond-scale
by: Hu, Tiancheng, et al.
Published: (2026)
by: Hu, Tiancheng, et al.
Published: (2026)
Limited Preemption in Real-Time Scheduling
by: Axel W. Krings
Published: (2000)
by: Axel W. Krings
Published: (2000)
Studying the Effect of Schedule Preemption on Dynamic Task Graph Scheduling
by: Khodabandehlou, Mohammadali, et al.
Published: (2026)
by: Khodabandehlou, Mohammadali, et al.
Published: (2026)
Learning-Augmented Online Scheduling with Parsimonious Preemption
by: Blue, Mugen, et al.
Published: (2026)
by: Blue, Mugen, et al.
Published: (2026)
LKD-KGC: Domain-Specific KG Construction via LLM-driven Knowledge Dependency Parsing
by: Sun, Jiaqi, et al.
Published: (2025)
by: Sun, Jiaqi, et al.
Published: (2025)
Opportunistic Scheduling for Optimal Spot Instance Savings in the Cloud
by: Bhuyan, Neelkamal, et al.
Published: (2026)
by: Bhuyan, Neelkamal, et al.
Published: (2026)
Priority Scheduling in the M/G/1 with Preemption Overhead
by: Ramakrishna, Shefali, et al.
Published: (2026)
by: Ramakrishna, Shefali, et al.
Published: (2026)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
by: Li, Bingyao, et al.
Published: (2024)
by: Li, Bingyao, et al.
Published: (2024)
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
Classification of secant defective manifolds near the extremal case
by: Han, Kangjin
Published: (2011)
by: Han, Kangjin
Published: (2011)
Towards Multi-Model LLM Schedulers: Empirical Insights into Offloading and Preemption
by: Yildiz, Mert, et al.
Published: (2026)
by: Yildiz, Mert, et al.
Published: (2026)
Validation of the GFS model for Gyrokinetic Stability of NSTX Pedestal Data
by: Yang, M., et al.
Published: (2025)
by: Yang, M., et al.
Published: (2025)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
by: Lettich, Francesco, et al.
Published: (2024)
by: Lettich, Francesco, et al.
Published: (2024)
SVFusion: A CPU-GPU Co-Processing Architecture for Large-Scale Real-Time Vector Search
by: Peng, Yuchen, et al.
Published: (2026)
by: Peng, Yuchen, et al.
Published: (2026)
Harmonic Decomposition in Data Sketches
by: Wang, Dingyu
Published: (2024)
by: Wang, Dingyu
Published: (2024)
Optimal Protocols for 2-Party Contention Resolution
by: Wang, Dingyu
Published: (2024)
by: Wang, Dingyu
Published: (2024)
Multi-dimensional Approximate Counting
by: Wang, Dingyu
Published: (2024)
by: Wang, Dingyu
Published: (2024)
Preemption Revisited: Multi-Threshold Preemption Policies for AoI Minimization
by: Liyanaarachchi, Sahan, et al.
Published: (2026)
by: Liyanaarachchi, Sahan, et al.
Published: (2026)
Limited Preemption of the 3-Phase Task Model using Preemption Thresholds
by: Thilakasiri, Thilanka, et al.
Published: (2025)
by: Thilakasiri, Thilanka, et al.
Published: (2025)
CAN GFS GLOBAL FORECASTS BE USEFUL FOR ASTRONOMERS?
by: J. C. Marín
Published: (2011)
by: J. C. Marín
Published: (2011)
Status Updating in Two-Way Delay Systems with Preemption
by: Yang, Jinxin, et al.
Published: (2026)
by: Yang, Jinxin, et al.
Published: (2026)
Particle-based Instance-aware Semantic Occupancy Mapping in Dynamic Environments
by: Chen, Gang, et al.
Published: (2024)
by: Chen, Gang, et al.
Published: (2024)
PAST: A Primary-Auxiliary Spatio-Temporal Network for Traffic Time Series Imputation
by: Hu, Hanwen, et al.
Published: (2025)
by: Hu, Hanwen, et al.
Published: (2025)
Simulating Dynamic Cloud Marketspaces: Modeling Spot Instance Behavior and Scheduling with CloudSim Plus
by: Goldgruber, Christoph, et al.
Published: (2025)
by: Goldgruber, Christoph, et al.
Published: (2025)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
by: Ting, Hsu-Tzu, et al.
Published: (2025)
by: Ting, Hsu-Tzu, et al.
Published: (2025)
Instance-aware Image Colorization with Controllable Textual Descriptions and Segmentation Masks
by: An, Yanru, et al.
Published: (2025)
by: An, Yanru, et al.
Published: (2025)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
Research on soil‐borne diseases of ginseng due to root exudates
by: Sen Jia, et al.
Published: (2024)
by: Sen Jia, et al.
Published: (2024)
EIMC: Efficient Instance-aware Multi-modal Collaborative Perception
by: Yang, Kang, et al.
Published: (2026)
by: Yang, Kang, et al.
Published: (2026)
Similar Items
-
MOPAR: A Model Partitioning Framework for Deep Learning Inference Services on Serverless Platforms
by: Duan, Jiaang, et al.
Published: (2024) -
EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training
by: Yi, Qingao, et al.
Published: (2025) -
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
by: Zhang, Kaixuan, et al.
Published: (2026) -
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
by: Yang, Dingyu, et al.
Published: (2026) -
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
by: Hua, Qin, et al.
Published: (2024)