Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xinyi, Zhao, Hanyu, Xiao, Wencong, Jia, Xianyan, Xu, Fei, Li, Yong, Lin, Wei, Liu, Fangming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Llumnix: Dynamic Scheduling for Large Language Model Serving
von: Sun, Biao, et al.
Veröffentlicht: (2024)
von: Sun, Biao, et al.
Veröffentlicht: (2024)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
von: Luo, Yizhou, et al.
Veröffentlicht: (2024)
Deep Back-Filling: a Split Window Technique for Deep Online Cluster Job Scheduling
von: Wang, Lingfei, et al.
Veröffentlicht: (2024)
von: Wang, Lingfei, et al.
Veröffentlicht: (2024)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
von: Zojer, Patrick, et al.
Veröffentlicht: (2026)
von: Zojer, Patrick, et al.
Veröffentlicht: (2026)
Aryl: An Elastic Cluster Scheduler for Deep Learning
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
von: Luo, Ziyue, et al.
Veröffentlicht: (2025)
Data-Locality-Aware Task Assignment and Scheduling for Distributed Job Executions
von: Zhao, Hailiang, et al.
Veröffentlicht: (2024)
von: Zhao, Hailiang, et al.
Veröffentlicht: (2024)
An Elastic Job Scheduler for HPC Applications on the Cloud
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2025)
Hamava: Fault-tolerant Reconfigurable Geo-Replication on Heterogeneous Clusters
von: Mane, Tejas, et al.
Veröffentlicht: (2024)
von: Mane, Tejas, et al.
Veröffentlicht: (2024)
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
von: Chen, Chen, et al.
Veröffentlicht: (2026)
von: Chen, Chen, et al.
Veröffentlicht: (2026)
Scalable HPC Job Scheduling and Resource Management in SST
von: Abdurahman, Abubeker, et al.
Veröffentlicht: (2025)
von: Abdurahman, Abubeker, et al.
Veröffentlicht: (2025)
SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
von: Amrizal, Muhammad Alfian, et al.
Veröffentlicht: (2025)
von: Amrizal, Muhammad Alfian, et al.
Veröffentlicht: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
Deep Reinforcement Learning for Job Scheduling and Resource Management in Cloud Computing: An Algorithm-Level Review
von: Gu, Yan, et al.
Veröffentlicht: (2025)
von: Gu, Yan, et al.
Veröffentlicht: (2025)
Adaptive Job Scheduling in Quantum Clouds Using Reinforcement Learning
von: Luo, Waylon, et al.
Veröffentlicht: (2025)
von: Luo, Waylon, et al.
Veröffentlicht: (2025)
Metronome: Efficient Scheduling for Periodic Traffic Jobs with Network and Priority Awareness
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
PHWSOA: A Pareto-based Hybrid Whale-Seagull Scheduling for Multi-Objective Tasks in Cloud Computing
von: Zhao, Zhi, et al.
Veröffentlicht: (2025)
von: Zhao, Zhi, et al.
Veröffentlicht: (2025)
Air-FedGA: A Grouping Asynchronous Federated Learning Mechanism Exploiting Over-the-air Computation
von: Ma, Qianpiao, et al.
Veröffentlicht: (2025)
von: Ma, Qianpiao, et al.
Veröffentlicht: (2025)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
Quantifying the Carbon Reduction of DAG Workloads: A Job Shop Scheduling Perspective
von: Bostandoost, Roozbeh, et al.
Veröffentlicht: (2025)
von: Bostandoost, Roozbeh, et al.
Veröffentlicht: (2025)
LiveR: Fine-Grained Elasticity via Live Reconfiguration for Model Training
von: Liu, Haoyuan, et al.
Veröffentlicht: (2026)
von: Liu, Haoyuan, et al.
Veröffentlicht: (2026)
Maple: A Multi-agent System for Portable Deep Learning across Clusters
von: Wu, Molang, et al.
Veröffentlicht: (2025)
von: Wu, Molang, et al.
Veröffentlicht: (2025)
Asymptotically Optimal Scheduling of Multiple Parallelizable Job Classes
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
von: Berg, Benjamin, et al.
Veröffentlicht: (2024)
Alternative Mixed Integer Linear Programming Optimization for Joint Job Scheduling and Data Allocation in Grid Computing
von: Feng, Shengyu, et al.
Veröffentlicht: (2025)
von: Feng, Shengyu, et al.
Veröffentlicht: (2025)
Reconfigurable Heterogeneous Quorum Systems
von: Li, Xiao, et al.
Veröffentlicht: (2023)
von: Li, Xiao, et al.
Veröffentlicht: (2023)
Megha: Decentralized Global Fair Scheduling for Federated Clusters
von: Thiyyakat, Meghana, et al.
Veröffentlicht: (2021)
von: Thiyyakat, Meghana, et al.
Veröffentlicht: (2021)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
GPU Cluster Scheduling for Network-Sensitive Deep Learning
von: Sharma, Aakash, et al.
Veröffentlicht: (2024)
von: Sharma, Aakash, et al.
Veröffentlicht: (2024)
Efficient Data Labeling and Optimal Device Scheduling in HWNs Using Clustered Federated Semi-Supervised Learning
von: Hamood, Moqbel, et al.
Veröffentlicht: (2024)
von: Hamood, Moqbel, et al.
Veröffentlicht: (2024)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
von: Lin, Bin, et al.
Veröffentlicht: (2024)
von: Lin, Bin, et al.
Veröffentlicht: (2024)
Learning-Based Approaches for Job Shop Scheduling Problems: A Review
von: Rihane, Karima, et al.
Veröffentlicht: (2025)
von: Rihane, Karima, et al.
Veröffentlicht: (2025)
A Taxonomy of Schedulers -- Operating Systems, Clusters and Big Data Frameworks
von: Sliwko, Leszek
Veröffentlicht: (2025)
von: Sliwko, Leszek
Veröffentlicht: (2025)
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
von: Hanafy, Walid A., et al.
Veröffentlicht: (2025)
von: Hanafy, Walid A., et al.
Veröffentlicht: (2025)
Exploiting Multicast for Accelerating Collective Communication
von: Xu, Chao, et al.
Veröffentlicht: (2026)
von: Xu, Chao, et al.
Veröffentlicht: (2026)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
A Deep Reinforcement Learning Approach for Cost Optimized Workflow Scheduling in Cloud Computing Environments
von: Jayanetti, Amanda, et al.
Veröffentlicht: (2024)
von: Jayanetti, Amanda, et al.
Veröffentlicht: (2024)
GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
von: Duan, Jiaang, et al.
Veröffentlicht: (2025)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
von: Jain, Rutwik, et al.
Veröffentlicht: (2024)
von: Jain, Rutwik, et al.
Veröffentlicht: (2024)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
von: Wei, Yufan, et al.
Veröffentlicht: (2025)
von: Wei, Yufan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Llumnix: Dynamic Scheduling for Large Language Model Serving
von: Sun, Biao, et al.
Veröffentlicht: (2024) -
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
von: Luo, Yizhou, et al.
Veröffentlicht: (2024) -
Deep Back-Filling: a Split Window Technique for Deep Online Cluster Job Scheduling
von: Wang, Lingfei, et al.
Veröffentlicht: (2024) -
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
von: Zojer, Patrick, et al.
Veröffentlicht: (2026) -
Aryl: An Elastic Cluster Scheduler for Deep Learning
von: Li, Jiamin, et al.
Veröffentlicht: (2022)