Resource Allocation and Workload Scheduling for Large-Scale Distributed Deep Learning: A Survey
Fuente:
arXiv
Guardado en:
| Autores principales: | Liang, Feng, Zhang, Zhen, Lu, Haifeng, Li, Chengming, Leung, Victor C. M., Guo, Yanyi, Hu, Xiping |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Communication-Efficient Large-Scale Distributed Deep Learning: A Comprehensive Survey
por: Liang, Feng, et al.
Publicado: (2024)
por: Liang, Feng, et al.
Publicado: (2024)
Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
por: Stavrinides, Georgios L., et al.
Publicado: (2025)
por: Stavrinides, Georgios L., et al.
Publicado: (2025)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
por: Luo, Ziyue, et al.
Publicado: (2025)
por: Luo, Ziyue, et al.
Publicado: (2025)
Learning to Schedule: A Supervised Learning Framework for Network-Aware Scheduling of Data-Intensive Workloads
por: Timilsina, Sankalpa, et al.
Publicado: (2025)
por: Timilsina, Sankalpa, et al.
Publicado: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
por: Tang, Lujie, et al.
Publicado: (2024)
por: Tang, Lujie, et al.
Publicado: (2024)
MSARS: A Meta-Learning and Reinforcement Learning Framework for SLO Resource Allocation and Adaptive Scaling for Microservices
por: Hu, Kan, et al.
Publicado: (2024)
por: Hu, Kan, et al.
Publicado: (2024)
iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
por: Hu, Yi-Xiang, et al.
Publicado: (2026)
por: Hu, Yi-Xiang, et al.
Publicado: (2026)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
por: Xu, Jingwei, et al.
Publicado: (2025)
por: Xu, Jingwei, et al.
Publicado: (2025)
Eventually-Consistent Federated Scheduling for Data Center Workloads
por: Thiyyakat, Meghana, et al.
Publicado: (2023)
por: Thiyyakat, Meghana, et al.
Publicado: (2023)
Deep Reinforcement Learning-based Methods for Resource Scheduling in Cloud Computing: A Review and Future Directions
por: Zhou, Guangyao, et al.
Publicado: (2021)
por: Zhou, Guangyao, et al.
Publicado: (2021)
Dynamic Client Clustering, Bandwidth Allocation, and Workload Optimization for Semi-synchronous Federated Learning
por: Yu, Liangkun, et al.
Publicado: (2024)
por: Yu, Liangkun, et al.
Publicado: (2024)
Sustainable Graph Analytics Workload Scheduling with Evolutionary Reinforcement Learning in Edge-Cloud Systems
por: Ramicetty, P., et al.
Publicado: (2026)
por: Ramicetty, P., et al.
Publicado: (2026)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
por: Zhu, Botao, et al.
Publicado: (2025)
por: Zhu, Botao, et al.
Publicado: (2025)
Decouple and Decompose: Scaling Resource Allocation with DeDe
por: Xu, Zhiying, et al.
Publicado: (2024)
por: Xu, Zhiying, et al.
Publicado: (2024)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
por: Svedas, Jonas, et al.
Publicado: (2025)
por: Svedas, Jonas, et al.
Publicado: (2025)
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
por: Chen, Chen, et al.
Publicado: (2026)
por: Chen, Chen, et al.
Publicado: (2026)
Task Scheduling in Geo-Distributed Computing: A Survey
por: Wu, Yujian, et al.
Publicado: (2025)
por: Wu, Yujian, et al.
Publicado: (2025)
Scheduling of Distributed Applications on the Computing Continuum: A Survey
por: Mehran, Narges, et al.
Publicado: (2024)
por: Mehran, Narges, et al.
Publicado: (2024)
AI Surrogate Model for Distributed Computing Workloads
por: Park, David K., et al.
Publicado: (2024)
por: Park, David K., et al.
Publicado: (2024)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
por: Wu, Xin, et al.
Publicado: (2026)
por: Wu, Xin, et al.
Publicado: (2026)
Distributed Hierarchical Machine Learning for Joint Resource Allocation and Slice Selection in In-Network Edge Systems
por: Rashid, Sulaiman Muhammad, et al.
Publicado: (2025)
por: Rashid, Sulaiman Muhammad, et al.
Publicado: (2025)
Duration-Informed Workload Scheduler
por: Loreti, Daniela, et al.
Publicado: (2026)
por: Loreti, Daniela, et al.
Publicado: (2026)
Quantifying the Carbon Reduction of DAG Workloads: A Job Shop Scheduling Perspective
por: Bostandoost, Roozbeh, et al.
Publicado: (2025)
por: Bostandoost, Roozbeh, et al.
Publicado: (2025)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
por: Zojer, Patrick, et al.
Publicado: (2026)
por: Zojer, Patrick, et al.
Publicado: (2026)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
por: Jain, Rutwik, et al.
Publicado: (2024)
por: Jain, Rutwik, et al.
Publicado: (2024)
Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
por: Zhao, Wei, et al.
Publicado: (2024)
por: Zhao, Wei, et al.
Publicado: (2024)
Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
por: Luo, Yizhou, et al.
Publicado: (2024)
por: Luo, Yizhou, et al.
Publicado: (2024)
Data Management System Analysis for Distributed Computing Workloads
por: Hsu, Kuan-Chieh, et al.
Publicado: (2025)
por: Hsu, Kuan-Chieh, et al.
Publicado: (2025)
Adaptive Resource Allocation for Workflow Containerization on Kubernetes
por: Shan, Chenggang, et al.
Publicado: (2023)
por: Shan, Chenggang, et al.
Publicado: (2023)
Workflow-Driven Modeling for the Compute Continuum: An Optimization Approach to Automated System and Workload Scheduling
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
por: Ting, Hsu-Tzu, et al.
Publicado: (2025)
por: Ting, Hsu-Tzu, et al.
Publicado: (2025)
A Review of Tools and Techniques for Optimization of Workload Mapping and Scheduling in Heterogeneous HPC System
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
por: Sharma, Aasish Kumar, et al.
Publicado: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
por: Adnan, Muhammad, et al.
Publicado: (2024)
por: Adnan, Muhammad, et al.
Publicado: (2024)
Workload Schedulers -- Genesis, Algorithms and Differences
por: Sliwko, Leszek, et al.
Publicado: (2025)
por: Sliwko, Leszek, et al.
Publicado: (2025)
Distributed Load Balancing with Workload-Dependent Service Rates
por: Zhang, Wenxin, et al.
Publicado: (2024)
por: Zhang, Wenxin, et al.
Publicado: (2024)
TD3-Sched: Learning to Orchestrate Container-based Cloud-Edge Resources via Distributed Reinforcement Learning
por: Song, Shengye, et al.
Publicado: (2025)
por: Song, Shengye, et al.
Publicado: (2025)
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
por: Iserte, Sergio, et al.
Publicado: (2025)
por: Iserte, Sergio, et al.
Publicado: (2025)
ARC-V: Vertical Resource Adaptivity for HPC Workloads in Containerized Environments
por: Medeiros, Daniel, et al.
Publicado: (2025)
por: Medeiros, Daniel, et al.
Publicado: (2025)
Workload Buoyancy: Keeping Apps Afloat by Identifying Shared Resource Bottlenecks
por: Larsson, Oliver, et al.
Publicado: (2026)
por: Larsson, Oliver, et al.
Publicado: (2026)
Adaptive, Efficient and Fair Resource Allocation in Cloud Datacenters leveraging Weighted A3C Deep Reinforcement Learning
por: Kumari, Suchi, et al.
Publicado: (2025)
por: Kumari, Suchi, et al.
Publicado: (2025)
Ejemplares similares
-
Communication-Efficient Large-Scale Distributed Deep Learning: A Comprehensive Survey
por: Liang, Feng, et al.
Publicado: (2024) -
Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
por: Stavrinides, Georgios L., et al.
Publicado: (2025) -
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
por: Luo, Ziyue, et al.
Publicado: (2025) -
Learning to Schedule: A Supervised Learning Framework for Network-Aware Scheduling of Data-Intensive Workloads
por: Timilsina, Sankalpa, et al.
Publicado: (2025) -
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
por: Tang, Lujie, et al.
Publicado: (2024)