Towards Cloud Efficiency with Large-scale Workload Characterization
Fuente:
arXiv
Saved in:
| Main Authors: | Parayil, Anjaly, Zhang, Jue, Qin, Xiaoting, Goiri, Íñigo, Huang, Lexiang, Zhu, Timothy, Bansal, Chetan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024)
by: Huang, Lexiang, et al.
Published: (2024)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)
by: Jain, Kunal, et al.
Published: (2024)
A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
by: Bastos, Anson, et al.
Published: (2026)
by: Bastos, Anson, et al.
Published: (2026)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
by: Biswas, Anish, et al.
Published: (2026)
by: Biswas, Anish, et al.
Published: (2026)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Cloud abstractions for AI workloads
by: Canini, Marco, et al.
Published: (2025)
by: Canini, Marco, et al.
Published: (2025)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Towards Resource-Efficient Compound AI Systems
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
by: Xiang, Yuxing, et al.
Published: (2025)
by: Xiang, Yuxing, et al.
Published: (2025)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
by: Qiu, Haoran, et al.
Published: (2025)
by: Qiu, Haoran, et al.
Published: (2025)
Dependency Aware Incident Linking in Large Cloud Systems
by: Ghosh, Supriyo, et al.
Published: (2024)
by: Ghosh, Supriyo, et al.
Published: (2024)
Orchestrating Mixed-Criticality Cloud Workloads in Reconfigurable Manufacturing Systems
by: Barletta, Marco, et al.
Published: (2024)
by: Barletta, Marco, et al.
Published: (2024)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
by: Chazapis, Antony, et al.
Published: (2024)
by: Chazapis, Antony, et al.
Published: (2024)
GreenFaaS: Maximizing Energy Efficiency of HPC Workloads with FaaS
by: Kamatar, Alok, et al.
Published: (2024)
by: Kamatar, Alok, et al.
Published: (2024)
An Empirical Study of Production Incidents in Generative AI Cloud Services
by: Yan, Haoran, et al.
Published: (2025)
by: Yan, Haoran, et al.
Published: (2025)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
by: Saurez, Enrique, et al.
Published: (2024)
by: Saurez, Enrique, et al.
Published: (2024)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Characterizing Production GPU Workloads using System-wide Telemetry Data
by: Cankur, Onur, et al.
Published: (2025)
by: Cankur, Onur, et al.
Published: (2025)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
Spatio-Temporal Shifting to Reduce Carbon, Water, and Land-Use Footprints of Cloud Workloads
by: Attenni, Giulio, et al.
Published: (2025)
by: Attenni, Giulio, et al.
Published: (2025)
Sustainable Graph Analytics Workload Scheduling with Evolutionary Reinforcement Learning in Edge-Cloud Systems
by: Ramicetty, P., et al.
Published: (2026)
by: Ramicetty, P., et al.
Published: (2026)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
by: Chen, Chang, et al.
Published: (2025)
by: Chen, Chang, et al.
Published: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
by: Tang, Lujie, et al.
Published: (2024)
by: Tang, Lujie, et al.
Published: (2024)
LLMSched: Uncertainty-Aware Workload Scheduling for Compound LLM Applications
by: Zhu, Botao, et al.
Published: (2025)
by: Zhu, Botao, et al.
Published: (2025)
A Framework for Carbon-aware Real-Time Workload Management in Clouds using Renewables-driven Cores
by: Hewage, Tharindu B., et al.
Published: (2024)
by: Hewage, Tharindu B., et al.
Published: (2024)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
by: Wen, Linfeng, et al.
Published: (2024)
by: Wen, Linfeng, et al.
Published: (2024)
Ksurf: Attention Kalman Filter and Principal Component Analysis for Prediction under Highly Variable Cloud Workloads
by: Dang'ana, Michael, et al.
Published: (2024)
by: Dang'ana, Michael, et al.
Published: (2024)
Towards the Distributed Large-scale k-NN Graph Construction by Graph Merge
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
by: Qiu, Haoran, et al.
Published: (2026)
by: Qiu, Haoran, et al.
Published: (2026)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
by: Du, Kuntai, et al.
Published: (2025)
by: Du, Kuntai, et al.
Published: (2025)
Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
by: Stavrinides, Georgios L., et al.
Published: (2025)
by: Stavrinides, Georgios L., et al.
Published: (2025)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
by: Wu, Xin, et al.
Published: (2026)
by: Wu, Xin, et al.
Published: (2026)
Agentic AI Workload Characteristics
by: Yuan, Yichao, et al.
Published: (2026)
by: Yuan, Yichao, et al.
Published: (2026)
BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
by: Duan, Tao, et al.
Published: (2025)
by: Duan, Tao, et al.
Published: (2025)
CloudFormer: An Attention-based Performance Prediction for Public Clouds with Unknown Workload
by: Shahbazinia, Amirhossein, et al.
Published: (2025)
by: Shahbazinia, Amirhossein, et al.
Published: (2025)
Distributed Load Balancing with Workload-Dependent Service Rates
by: Zhang, Wenxin, et al.
Published: (2024)
by: Zhang, Wenxin, et al.
Published: (2024)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
Accelerating Compound LLM Training Workloads with Maestro
by: Yuan, Xiulong, et al.
Published: (2026)
by: Yuan, Xiulong, et al.
Published: (2026)
Similar Items
-
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024) -
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024) -
A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
by: Bastos, Anson, et al.
Published: (2026) -
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
by: Biswas, Anish, et al.
Published: (2026) -
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)