Workload Intelligence: Punching Holes Through the Cloud Abstraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Lexiang, Parayil, Anjaly, Zhang, Jue, Qin, Xiaoting, Bansal, Chetan, Stojkovic, Jovan, Zardoshti, Pantea, Misra, Pulkit, Cortez, Eli, Ghelman, Raphael, Goiri, Íñigo, Rajmohan, Saravan, Kleewein, Jim, Fonseca, Rodrigo, Zhu, Timothy, Bianchini, Ricardo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Cloud Efficiency with Large-scale Workload Characterization
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
von: Jain, Kunal, et al.
Veröffentlicht: (2024)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
Intelligent Monitoring Framework for Cloud Services: A Data-Driven Approach
von: Srinivas, Pooja, et al.
Veröffentlicht: (2024)
von: Srinivas, Pooja, et al.
Veröffentlicht: (2024)
Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems
von: Hussain, Fiza, et al.
Veröffentlicht: (2025)
von: Hussain, Fiza, et al.
Veröffentlicht: (2025)
X-lifecycle Learning for Cloud Incident Management using LLMs
von: Goel, Drishti, et al.
Veröffentlicht: (2024)
von: Goel, Drishti, et al.
Veröffentlicht: (2024)
Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud Platforms
von: Reidys, Benjamin, et al.
Veröffentlicht: (2025)
von: Reidys, Benjamin, et al.
Veröffentlicht: (2025)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2024)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
von: Bastos, Anson, et al.
Veröffentlicht: (2026)
von: Bastos, Anson, et al.
Veröffentlicht: (2026)
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
von: Xia, Menglin, et al.
Veröffentlicht: (2026)
von: Xia, Menglin, et al.
Veröffentlicht: (2026)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2025)
REFA: Reference Free Alignment for multi-preference optimization
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
Risk-aware Adaptive Virtual CPU Oversubscription in Microsoft Cloud via Prototypical Human-in-the-loop Imitation Learning
von: Wang, Lu, et al.
Veröffentlicht: (2024)
von: Wang, Lu, et al.
Veröffentlicht: (2024)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
von: Biswas, Anish, et al.
Veröffentlicht: (2026)
AutoAdapt: An Automated Domain Adaptation Framework for LLMs
von: Sinha, Sidharth, et al.
Veröffentlicht: (2026)
von: Sinha, Sidharth, et al.
Veröffentlicht: (2026)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
von: Zhang, Jue, et al.
Veröffentlicht: (2025)
von: Zhang, Jue, et al.
Veröffentlicht: (2025)
Dependency Aware Incident Linking in Large Cloud Systems
von: Ghosh, Supriyo, et al.
Veröffentlicht: (2024)
von: Ghosh, Supriyo, et al.
Veröffentlicht: (2024)
Automated Root Causing of Cloud Incidents using In-Context Learning with GPT-4
von: Zhang, Xuchao, et al.
Veröffentlicht: (2024)
von: Zhang, Xuchao, et al.
Veröffentlicht: (2024)
Exploring LLM-based Agents for Root Cause Analysis
von: Roy, Devjeet, et al.
Veröffentlicht: (2024)
von: Roy, Devjeet, et al.
Veröffentlicht: (2024)
Ensuring Fair LLM Serving Amid Diverse Applications
von: Khan, Redwan Ibne Seraj, et al.
Veröffentlicht: (2024)
von: Khan, Redwan Ibne Seraj, et al.
Veröffentlicht: (2024)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
von: Wang, Shouju, et al.
Veröffentlicht: (2025)
von: Wang, Shouju, et al.
Veröffentlicht: (2025)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
CARMO: Dynamic Criteria Generation for Context-Aware Reward Modelling
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
EvidenT: An Evidence-Preserving Framework for Iterative System-Level Package Repair
von: Zhao, Chenyu, et al.
Veröffentlicht: (2026)
von: Zhao, Chenyu, et al.
Veröffentlicht: (2026)
eARCO: Efficient Automated Root Cause Analysis with Prompt Optimization
von: Goel, Drishti, et al.
Veröffentlicht: (2025)
von: Goel, Drishti, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Compound AI Systems
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
von: Chaudhry, Gohar Irfan, et al.
Veröffentlicht: (2025)
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC
von: Soni, Aditya, et al.
Veröffentlicht: (2024)
von: Soni, Aditya, et al.
Veröffentlicht: (2024)
Navigating the Unknown: A Chat-Based Collaborative Interface for Personalized Exploratory Tasks
von: Peng, Yingzhe, et al.
Veröffentlicht: (2024)
von: Peng, Yingzhe, et al.
Veröffentlicht: (2024)
The Vision of Autonomic Computing: Can LLMs Make It a Reality?
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2024)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
von: Tan, Rongyuan, et al.
Veröffentlicht: (2026)
von: Tan, Rongyuan, et al.
Veröffentlicht: (2026)
Large Language Models can Deliver Accurate and Interpretable Time Series Anomaly Detection
von: Liu, Jun, et al.
Veröffentlicht: (2024)
von: Liu, Jun, et al.
Veröffentlicht: (2024)
AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
von: Chen, Yinfang, et al.
Veröffentlicht: (2025)
von: Chen, Yinfang, et al.
Veröffentlicht: (2025)
CONOCIMIENTO DE LAS PRÁCTICAS DE AUTOCUIDADO EN LOS PIES DE LOS INDIVIDUOS CON DIABETES MELLITUS ATENDIDOS EN UNA UNIDAD BÁSICA DE SALUD.
von: L. Gack Ghelman
Veröffentlicht: (2009)
von: L. Gack Ghelman
Veröffentlicht: (2009)
Cost-Aware Retrieval-Augmentation Reasoning Models with Adaptive Retrieval Depth
von: Hashemi, Helia, et al.
Veröffentlicht: (2025)
von: Hashemi, Helia, et al.
Veröffentlicht: (2025)
Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
von: Jaiswal, Shashwat, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Cloud Efficiency with Large-scale Workload Characterization
von: Parayil, Anjaly, et al.
Veröffentlicht: (2024) -
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
von: Jain, Kunal, et al.
Veröffentlicht: (2024) -
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025) -
Intelligent Monitoring Framework for Cloud Services: A Data-Driven Approach
von: Srinivas, Pooja, et al.
Veröffentlicht: (2024) -
Attention Enhanced Entity Recommendation for Intelligent Monitoring in Cloud Systems
von: Hussain, Fiza, et al.
Veröffentlicht: (2025)