An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Hang, Zhu, Liwen, Shan, Zhao, Qiao, Bo, Yang, Fangkai, Qin, Si, Luo, Chuan, Lin, Qingwei, Yang, Yuwen, Virdi, Gurpreet, Rajmohan, Saravan, Zhang, Dongmei, Moscibroda, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
by: Liao, Mengqi, et al.
Published: (2026)
by: Liao, Mengqi, et al.
Published: (2026)
An Empirical Study of Production Incidents in Generative AI Cloud Services
by: Yan, Haoran, et al.
Published: (2025)
by: Yan, Haoran, et al.
Published: (2025)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024)
by: Huang, Lexiang, et al.
Published: (2024)
The Vision of Autonomic Computing: Can LLMs Make It a Reality?
by: Zhang, Zhiyang, et al.
Published: (2024)
by: Zhang, Zhiyang, et al.
Published: (2024)
UFO3: Weaving the Digital Agent Galaxy
by: Zhang, Chaoyun, et al.
Published: (2025)
by: Zhang, Chaoyun, et al.
Published: (2025)
Sustainable Graph Analytics Workload Scheduling with Evolutionary Reinforcement Learning in Edge-Cloud Systems
by: Ramicetty, P., et al.
Published: (2026)
by: Ramicetty, P., et al.
Published: (2026)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)
by: Jain, Kunal, et al.
Published: (2024)
Dependency Aware Incident Linking in Large Cloud Systems
by: Ghosh, Supriyo, et al.
Published: (2024)
by: Ghosh, Supriyo, et al.
Published: (2024)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
by: Ting, Hsu-Tzu, et al.
Published: (2025)
by: Ting, Hsu-Tzu, et al.
Published: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
by: Tang, Lujie, et al.
Published: (2024)
by: Tang, Lujie, et al.
Published: (2024)
KCES: A Workflow Containerization Scheduling Scheme Under Cloud-Edge Collaboration Framework
by: Shan, Chenggang, et al.
Published: (2024)
by: Shan, Chenggang, et al.
Published: (2024)
Eventually-Consistent Federated Scheduling for Data Center Workloads
by: Thiyyakat, Meghana, et al.
Published: (2023)
by: Thiyyakat, Meghana, et al.
Published: (2023)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
by: Zhang, Ping, et al.
Published: (2024)
by: Zhang, Ping, et al.
Published: (2024)
Prediction-Assisted Online Distributed Deep Learning Workload Scheduling in GPU Clusters
by: Luo, Ziyue, et al.
Published: (2025)
by: Luo, Ziyue, et al.
Published: (2025)
Why does Prediction Accuracy Decrease over Time? Uncertain Positive Learning for Cloud Failure Prediction
by: Li, Haozhe, et al.
Published: (2024)
by: Li, Haozhe, et al.
Published: (2024)
Duration-Informed Workload Scheduler
by: Loreti, Daniela, et al.
Published: (2026)
by: Loreti, Daniela, et al.
Published: (2026)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
by: Zhu, Botao, et al.
Published: (2025)
by: Zhu, Botao, et al.
Published: (2025)
Learning to Schedule: A Supervised Learning Framework for Network-Aware Scheduling of Data-Intensive Workloads
by: Timilsina, Sankalpa, et al.
Published: (2025)
by: Timilsina, Sankalpa, et al.
Published: (2025)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Workload Schedulers -- Genesis, Algorithms and Differences
by: Sliwko, Leszek, et al.
Published: (2025)
by: Sliwko, Leszek, et al.
Published: (2025)
PHWSOA: A Pareto-based Hybrid Whale-Seagull Scheduling for Multi-Objective Tasks in Cloud Computing
by: Zhao, Zhi, et al.
Published: (2025)
by: Zhao, Zhi, et al.
Published: (2025)
Orchestrating Mixed-Criticality Cloud Workloads in Reconfigurable Manufacturing Systems
by: Barletta, Marco, et al.
Published: (2024)
by: Barletta, Marco, et al.
Published: (2024)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
by: Chazapis, Antony, et al.
Published: (2024)
by: Chazapis, Antony, et al.
Published: (2024)
Quantifying the Carbon Reduction of DAG Workloads: A Job Shop Scheduling Perspective
by: Bostandoost, Roozbeh, et al.
Published: (2025)
by: Bostandoost, Roozbeh, et al.
Published: (2025)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
by: Zojer, Patrick, et al.
Published: (2026)
by: Zojer, Patrick, et al.
Published: (2026)
Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
by: Stavrinides, Georgios L., et al.
Published: (2025)
by: Stavrinides, Georgios L., et al.
Published: (2025)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
by: Jain, Rutwik, et al.
Published: (2024)
by: Jain, Rutwik, et al.
Published: (2024)
Competitive Capacitated Online Recoloring
by: Rajaraman, Rajmohan, et al.
Published: (2024)
by: Rajaraman, Rajmohan, et al.
Published: (2024)
Workflow-Driven Modeling for the Compute Continuum: An Optimization Approach to Automated System and Workload Scheduling
by: Sharma, Aasish Kumar, et al.
Published: (2025)
by: Sharma, Aasish Kumar, et al.
Published: (2025)
A Review of Tools and Techniques for Optimization of Workload Mapping and Scheduling in Heterogeneous HPC System
by: Sharma, Aasish Kumar, et al.
Published: (2025)
by: Sharma, Aasish Kumar, et al.
Published: (2025)
A Deep Reinforcement Learning Approach for Cost Optimized Workflow Scheduling in Cloud Computing Environments
by: Jayanetti, Amanda, et al.
Published: (2024)
by: Jayanetti, Amanda, et al.
Published: (2024)
Hestia: Hyperthread-Level Scheduling for Cloud Microservices with Interference-Aware Attention
by: Yang, Dingyu, et al.
Published: (2026)
by: Yang, Dingyu, et al.
Published: (2026)
Deoxys: A Causal Inference Engine for Unhealthy Node Mitigation in Large-scale Cloud Infrastructure
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
Deep Reinforcement Learning-based Methods for Resource Scheduling in Cloud Computing: A Review and Future Directions
by: Zhou, Guangyao, et al.
Published: (2021)
by: Zhou, Guangyao, et al.
Published: (2021)
Spatio-Temporal Shifting to Reduce Carbon, Water, and Land-Use Footprints of Cloud Workloads
by: Attenni, Giulio, et al.
Published: (2025)
by: Attenni, Giulio, et al.
Published: (2025)
Adaptive Job Scheduling in Quantum Clouds Using Reinforcement Learning
by: Luo, Waylon, et al.
Published: (2025)
by: Luo, Waylon, et al.
Published: (2025)
An Elastic Job Scheduler for HPC Applications on the Cloud
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Efficient Probabilistic Workflow Scheduling for IaaS Clouds
by: Russo, Gabriele Russo, et al.
Published: (2024)
by: Russo, Gabriele Russo, et al.
Published: (2024)
Reinforcement Learning based Workflow Scheduling in Cloud and Edge Computing Environments: A Taxonomy, Review and Future Directions
by: Jayanetti, Amanda, et al.
Published: (2024)
by: Jayanetti, Amanda, et al.
Published: (2024)
Similar Items
-
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
by: Liao, Mengqi, et al.
Published: (2026) -
An Empirical Study of Production Incidents in Generative AI Cloud Services
by: Yan, Haoran, et al.
Published: (2025) -
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024) -
The Vision of Autonomic Computing: Can LLMs Make It a Reality?
by: Zhang, Zhiyang, et al.
Published: (2024) -
UFO3: Weaving the Digital Agent Galaxy
by: Zhang, Chaoyun, et al.
Published: (2025)