Aryl: An Elastic Cluster Scheduler for Deep Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Jiamin, Xu, Hong, Zhu, Yibo, Liu, Zherui, Guo, Chuanxiong, Wang, Cong |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025)
by: Luo, Ziyue, et al.
Published: (2025)
Accelerating Distributed MoE Training and Inference with Lina
by: Li, Jiamin, et al.
Published: (2022)
by: Li, Jiamin, et al.
Published: (2022)
Learning to Schedule Online Tasks with Bandit Feedback
by: Xu, Yongxin, et al.
Published: (2024)
by: Xu, Yongxin, et al.
Published: (2024)
GPU Cluster Scheduling for Network-Sensitive Deep Learning
by: Sharma, Aakash, et al.
Published: (2024)
by: Sharma, Aakash, et al.
Published: (2024)
Scheduling for On-Board Federated Learning with Satellite Clusters
by: Razmi, Nasrin, et al.
Published: (2024)
by: Razmi, Nasrin, et al.
Published: (2024)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
by: Wu, Bingyang, et al.
Published: (2025)
by: Wu, Bingyang, et al.
Published: (2025)
Semantic-Aware Scheduling for GPU Clusters with Large Language Models
by: Wang, Zerui, et al.
Published: (2025)
by: Wang, Zerui, et al.
Published: (2025)
Prioritizing Modalities: Flexible Importance Scheduling in Federated Multimodal Learning
by: Bian, Jieming, et al.
Published: (2024)
by: Bian, Jieming, et al.
Published: (2024)
Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling
by: Wang, Yujie, et al.
Published: (2024)
by: Wang, Yujie, et al.
Published: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Reinforcement Learning for Adaptive Resource Scheduling in Complex System Environments
by: Li, Pochun, et al.
Published: (2024)
by: Li, Pochun, et al.
Published: (2024)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
by: Zhong, Yinmin, et al.
Published: (2025)
by: Zhong, Yinmin, et al.
Published: (2025)
ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
by: Liu, Zedong, et al.
Published: (2025)
by: Liu, Zedong, et al.
Published: (2025)
Rubick: Exploiting Job Reconfigurability for Deep Learning Cluster Scheduling
by: Zhang, Xinyi, et al.
Published: (2024)
by: Zhang, Xinyi, et al.
Published: (2024)
Near-Optimal Resilient Aggregation Rules for Distributed Learning Using 1-Center and 1-Mean Clustering with Outliers
by: Yi, Yuhao, et al.
Published: (2023)
by: Yi, Yuhao, et al.
Published: (2023)
Federated Temporal Graph Clustering
by: Zhou, Zihao, et al.
Published: (2024)
by: Zhou, Zihao, et al.
Published: (2024)
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
by: Wang, Yuanqing, et al.
Published: (2026)
by: Wang, Yuanqing, et al.
Published: (2026)
Accelerating AIGC Services with Latent Action Diffusion Scheduling in Edge Networks
by: Xu, Changfu, et al.
Published: (2024)
by: Xu, Changfu, et al.
Published: (2024)
CSAFL: A Clustered Semi-Asynchronous Federated Learning Framework
by: Zhang, Yu, et al.
Published: (2021)
by: Zhang, Yu, et al.
Published: (2021)
Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
StraightLine: An End-to-End Resource-Aware Scheduler for Machine Learning Application Requests
by: Ching, Cheng-Wei, et al.
Published: (2024)
by: Ching, Cheng-Wei, et al.
Published: (2024)
DASH: Deterministic Attention Scheduling for High-throughput Reproducible LLM Training
by: Qiang, Xinwei, et al.
Published: (2026)
by: Qiang, Xinwei, et al.
Published: (2026)
Communication-Efficient Device Scheduling for Federated Learning Using Lyapunov Optimization
by: Perazzone, Jake B., et al.
Published: (2025)
by: Perazzone, Jake B., et al.
Published: (2025)
An Efficient Gradient-Aware Error-Bounded Lossy Compressor for Federated Learning
by: Ye, Zhijing, et al.
Published: (2025)
by: Ye, Zhijing, et al.
Published: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
by: Wu, Yongji, et al.
Published: (2024)
by: Wu, Yongji, et al.
Published: (2024)
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
Accelerating Hybrid Federated Learning Convergence under Partial Participation
by: Bian, Jieming, et al.
Published: (2023)
by: Bian, Jieming, et al.
Published: (2023)
Improving the Efficiency of a Deep Reinforcement Learning-Based Power Management System for HPC Clusters Using Curriculum Learning
by: Budiarjo, Thomas, et al.
Published: (2025)
by: Budiarjo, Thomas, et al.
Published: (2025)
Averaging Rate Scheduler for Decentralized Learning on Heterogeneous Data
by: Aketi, Sai Aparna, et al.
Published: (2024)
by: Aketi, Sai Aparna, et al.
Published: (2024)
Locality-aware Fair Scheduling in LLM Serving
by: Cao, Shiyi, et al.
Published: (2025)
by: Cao, Shiyi, et al.
Published: (2025)
FedClust: Optimizing Federated Learning on Non-IID Data through Weight-Driven Client Clustering
by: Islam, Md Sirajul, et al.
Published: (2024)
by: Islam, Md Sirajul, et al.
Published: (2024)
FLMarket: Enabling Privacy-preserved Pre-training Data Pricing for Federated Learning
by: Wen, Zhenyu, et al.
Published: (2024)
by: Wen, Zhenyu, et al.
Published: (2024)
Unleashing the Power of Continual Learning on Non-Centralized Devices: A Survey
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
Intelligent Task Scheduling for Microservices via A3C-Based Reinforcement Learning
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Deadline-Aware Online Scheduling for LLM Fine-Tuning with Spot Market Predictions
by: Kong, Linggao, et al.
Published: (2025)
by: Kong, Linggao, et al.
Published: (2025)
Device Scheduling and Assignment in Hierarchical Federated Learning for Internet of Things
by: Zhang, Tinghao, et al.
Published: (2024)
by: Zhang, Tinghao, et al.
Published: (2024)
DeFRiS: Silo-Cooperative IoT Applications Scheduling via Decentralized Federated Reinforcement Learning
by: Wang, Zhiyu, et al.
Published: (2026)
by: Wang, Zhiyu, et al.
Published: (2026)
SAFL: Structure-Aware Personalized Federated Learning via Client-Specific Clustering and SCSI-Guided Model Pruning
by: Li, Nan, et al.
Published: (2025)
by: Li, Nan, et al.
Published: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
by: Wu, Bingyang, et al.
Published: (2024)
by: Wu, Bingyang, et al.
Published: (2024)
Similar Items
-
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025) -
Accelerating Distributed MoE Training and Inference with Lina
by: Li, Jiamin, et al.
Published: (2022) -
Learning to Schedule Online Tasks with Bandit Feedback
by: Xu, Yongxin, et al.
Published: (2024) -
GPU Cluster Scheduling for Network-Sensitive Deep Learning
by: Sharma, Aakash, et al.
Published: (2024) -
Scheduling for On-Board Federated Learning with Satellite Clusters
by: Razmi, Nasrin, et al.
Published: (2024)