Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud Platforms
Fuente:
arXiv
Saved in:
| Main Authors: | Reidys, Benjamin, Zardoshti, Pantea, Goiri, Íñigo, Irvene, Celine, Berger, Daniel S., Ma, Haoran, Arya, Kapil, Cortez, Eli, Stark, Taylor, Bak, Eugene, Iyigun, Mehmet, Novaković, Stanko, Hsu, Lisa, Trueba, Karel, Pan, Abhisek, Bansal, Chetan, Rajmohan, Saravan, Huang, Jian, Bianchini, Ricardo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Risk-aware Adaptive Virtual CPU Oversubscription in Microsoft Cloud via Prototypical Human-in-the-loop Imitation Learning
by: Wang, Lu, et al.
Published: (2024)
by: Wang, Lu, et al.
Published: (2024)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024)
by: Huang, Lexiang, et al.
Published: (2024)
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
by: Wang, Lu, et al.
Published: (2024)
by: Wang, Lu, et al.
Published: (2024)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025)
by: Gupta, Taneesh, et al.
Published: (2025)
REFA: Reference Free Alignment for multi-preference optimization
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Intelligent Monitoring Framework for Cloud Services: A Data-Driven Approach
by: Srinivas, Pooja, et al.
Published: (2024)
by: Srinivas, Pooja, et al.
Published: (2024)
AutoAdapt: An Automated Domain Adaptation Framework for LLMs
by: Sinha, Sidharth, et al.
Published: (2026)
by: Sinha, Sidharth, et al.
Published: (2026)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)
by: Jain, Kunal, et al.
Published: (2024)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems
by: Hussain, Fiza, et al.
Published: (2025)
by: Hussain, Fiza, et al.
Published: (2025)
Dependency Aware Incident Linking in Large Cloud Systems
by: Ghosh, Supriyo, et al.
Published: (2024)
by: Ghosh, Supriyo, et al.
Published: (2024)
Automated Root Causing of Cloud Incidents using In-Context Learning with GPT-4
by: Zhang, Xuchao, et al.
Published: (2024)
by: Zhang, Xuchao, et al.
Published: (2024)
Exploring LLM-based Agents for Root Cause Analysis
by: Roy, Devjeet, et al.
Published: (2024)
by: Roy, Devjeet, et al.
Published: (2024)
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
by: Gupta, Taneesh, et al.
Published: (2024)
by: Gupta, Taneesh, et al.
Published: (2024)
EvidenT: An Evidence-Preserving Framework for Iterative System-Level Package Repair
by: Zhao, Chenyu, et al.
Published: (2026)
by: Zhao, Chenyu, et al.
Published: (2026)
eARCO: Efficient Automated Root Cause Analysis with Prompt Optimization
by: Goel, Drishti, et al.
Published: (2025)
by: Goel, Drishti, et al.
Published: (2025)
X-lifecycle Learning for Cloud Incident Management using LLMs
by: Goel, Drishti, et al.
Published: (2024)
by: Goel, Drishti, et al.
Published: (2024)
Intent-based System Design and Operation
by: Anand, Vaastav, et al.
Published: (2025)
by: Anand, Vaastav, et al.
Published: (2025)
Towards Resource-Efficient Compound AI Systems
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
Large Language Models can Deliver Accurate and Interpretable Time Series Anomaly Detection
by: Liu, Jun, et al.
Published: (2024)
by: Liu, Jun, et al.
Published: (2024)
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
by: Xia, Menglin, et al.
Published: (2026)
by: Xia, Menglin, et al.
Published: (2026)
AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
by: Chen, Yinfang, et al.
Published: (2025)
by: Chen, Yinfang, et al.
Published: (2025)
Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
by: Hashemi, Helia, et al.
Published: (2025)
by: Hashemi, Helia, et al.
Published: (2025)
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023)
by: Patel, Pratyush, et al.
Published: (2023)
A Rare Catastrophe: Three Cases of Aortic Root Dehiscence after Surgery
by: Taner İyigün
Published: (2023)
by: Taner İyigün
Published: (2023)
The Predictive Effects of Clinical Hematological Changes on Saphenous Graft Patency after Coronary Artery Surgery
by: Taner İyigün
Published: (2019)
by: Taner İyigün
Published: (2019)
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
by: Qiu, Haoran, et al.
Published: (2026)
by: Qiu, Haoran, et al.
Published: (2026)
Cloud abstractions for AI workloads
by: Canini, Marco, et al.
Published: (2025)
by: Canini, Marco, et al.
Published: (2025)
Synergistic Weak-Strong Collaboration by Aligning Preferences
by: Jiao, Yizhu, et al.
Published: (2025)
by: Jiao, Yizhu, et al.
Published: (2025)
Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents
by: Zhao, Chenyu, et al.
Published: (2026)
by: Zhao, Chenyu, et al.
Published: (2026)
Octopus: Enhancing CXL Memory Pods via Sparse Topology
by: Zhong, Yuhong, et al.
Published: (2025)
by: Zhong, Yuhong, et al.
Published: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Continuous Benchmark Generation for Evaluating Enterprise-scale LLM Agents
by: Saxena, Divyanshu, et al.
Published: (2025)
by: Saxena, Divyanshu, et al.
Published: (2025)
My CXL Pool Obviates Your PCIe Switch
by: Zhong, Yuhong, et al.
Published: (2025)
by: Zhong, Yuhong, et al.
Published: (2025)
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
by: Zhang, Jue, et al.
Published: (2025)
by: Zhang, Jue, et al.
Published: (2025)
Minerva: A Programmable Memory Test Benchmark for Language Models
by: Xia, Menglin, et al.
Published: (2025)
by: Xia, Menglin, et al.
Published: (2025)
Can Language Models Go Beyond Coding? Assessing the Capability of Language Models to Build Real-World Systems
by: Zhao, Chenyu, et al.
Published: (2025)
by: Zhao, Chenyu, et al.
Published: (2025)
Similar Items
-
Risk-aware Adaptive Virtual CPU Oversubscription in Microsoft Cloud via Prototypical Human-in-the-loop Imitation Learning
by: Wang, Lu, et al.
Published: (2024) -
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024) -
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
by: Wang, Lu, et al.
Published: (2024) -
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
by: Gupta, Taneesh, et al.
Published: (2025) -
REFA: Reference Free Alignment for multi-preference optimization
by: Gupta, Taneesh, et al.
Published: (2024)