Semantic-Aware Scheduling for GPU Clusters with Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zerui, Hu, Qinghao, Klimovic, Ana, Zhang, Tianwei, Wen, Yonggang, Sun, Peng, Lin, Dahua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Characterization of Large Language Model Development in the Datacenter
von: Hu, Qinghao, et al.
Veröffentlicht: (2024)
von: Hu, Qinghao, et al.
Veröffentlicht: (2024)
DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
von: Yao, Xiaozhe, et al.
Veröffentlicht: (2023)
von: Yao, Xiaozhe, et al.
Veröffentlicht: (2023)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
TorchGT: A Holistic System for Large-scale Graph Transformer Training
von: Zhang, Meng, et al.
Veröffentlicht: (2024)
von: Zhang, Meng, et al.
Veröffentlicht: (2024)
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
Understanding GPU Resource Interference One Level Deeper
von: Elvinger, Paul, et al.
Veröffentlicht: (2025)
von: Elvinger, Paul, et al.
Veröffentlicht: (2025)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2024)
von: Jain, Rutwik, et al.
Veröffentlicht: (2024)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
von: Chen, Chang, et al.
Veröffentlicht: (2025)
von: Chen, Chang, et al.
Veröffentlicht: (2025)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
GPU Cluster Scheduling for Network-Sensitive Deep Learning
von: Sharma, Aakash, et al.
Veröffentlicht: (2024)
von: Sharma, Aakash, et al.
Veröffentlicht: (2024)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
Llumnix: Dynamic Scheduling for Large Language Model Serving
von: Sun, Biao, et al.
Veröffentlicht: (2024)
von: Sun, Biao, et al.
Veröffentlicht: (2024)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules
von: Pan, Xinglin, et al.
Veröffentlicht: (2024)
von: Pan, Xinglin, et al.
Veröffentlicht: (2024)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
von: Hu, Qinghao, et al.
Veröffentlicht: (2025)
von: Hu, Qinghao, et al.
Veröffentlicht: (2025)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
RL in the Wild: Characterizing RLVR Training in LLM Deployment
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2025)
LLMPerf: GPU Performance Modeling meets Large Language Models
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
von: Nguyen, Khoi N. M., et al.
Veröffentlicht: (2025)
BandPilot: Towards Performance- and Contention-Aware GPU Dispatching in AI Clusters
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
von: Zhang, Kunming, et al.
Veröffentlicht: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
von: Zhang, Mingjun, et al.
Veröffentlicht: (2025)
von: Zhang, Mingjun, et al.
Veröffentlicht: (2025)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
von: Ting, Hsu-Tzu, et al.
Veröffentlicht: (2025)
GCAPS: GPU Context-Aware Preemptive Priority-based Scheduling for Real-Time Tasks
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
von: Wang, Yidi, et al.
Veröffentlicht: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2024)
tf.data service: A Case for Disaggregating ML Input Data Processing
von: Audibert, Andrew, et al.
Veröffentlicht: (2022)
von: Audibert, Andrew, et al.
Veröffentlicht: (2022)
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
von: Xiao, Youshao, et al.
Veröffentlicht: (2024)
von: Xiao, Youshao, et al.
Veröffentlicht: (2024)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
Hybrid Learning and Optimization-Based Dynamic Scheduling for DL Workloads on Heterogeneous GPU Clusters
von: Dongare, Shruti, et al.
Veröffentlicht: (2025)
von: Dongare, Shruti, et al.
Veröffentlicht: (2025)
Aryl: An Elastic Cluster Scheduler for Deep Learning
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
Scheduling for On-Board Federated Learning with Satellite Clusters
von: Razmi, Nasrin, et al.
Veröffentlicht: (2024)
von: Razmi, Nasrin, et al.
Veröffentlicht: (2024)
FIKIT: Priority-Based Real-time GPU Multi-tasking Scheduling with Kernel Identification
von: Wu, Wenqing
Veröffentlicht: (2023)
von: Wu, Wenqing
Veröffentlicht: (2023)
Ähnliche Einträge
-
Characterization of Large Language Model Development in the Datacenter
von: Hu, Qinghao, et al.
Veröffentlicht: (2024) -
DeltaZip: Efficient Serving of Multiple Full-Model-Tuned LLMs
von: Yao, Xiaozhe, et al.
Veröffentlicht: (2023) -
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024) -
TorchGT: A Holistic System for Large-scale Graph Transformer Training
von: Zhang, Meng, et al.
Veröffentlicht: (2024) -
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)