An Empirical Study of Production Incidents in Generative AI Cloud Services
Fuente:
arXiv
Guardado en:
| Autores principales: | Yan, Haoran, Chen, Yinfang, Ma, Minghua, Wen, Ming, Lu, Shan, Zhang, Shenglin, Xu, Tianyin, Wang, Rujia, Bansal, Chetan, Rajmohan, Saravan, Lin, Qingwei, Zhang, Chaoyun, Zhang, Dongmei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
por: Chen, Yinfang, et al.
Publicado: (2025)
por: Chen, Yinfang, et al.
Publicado: (2025)
Dependency Aware Incident Linking in Large Cloud Systems
por: Ghosh, Supriyo, et al.
Publicado: (2024)
por: Ghosh, Supriyo, et al.
Publicado: (2024)
UFO3: Weaving the Digital Agent Galaxy
por: Zhang, Chaoyun, et al.
Publicado: (2025)
por: Zhang, Chaoyun, et al.
Publicado: (2025)
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
por: Liao, Mengqi, et al.
Publicado: (2026)
por: Liao, Mengqi, et al.
Publicado: (2026)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
por: Jaiswal, Shashwat, et al.
Publicado: (2025)
A Holistic Framework for Automated Configuration Recommendation for Cloud Service Monitoring
por: Bastos, Anson, et al.
Publicado: (2026)
por: Bastos, Anson, et al.
Publicado: (2026)
Building AI Agents for Autonomous Clouds: Challenges and Design Principles
por: Shetty, Manish, et al.
Publicado: (2024)
por: Shetty, Manish, et al.
Publicado: (2024)
Deoxys: A Causal Inference Engine for Unhealthy Node Mitigation in Large-scale Cloud Infrastructure
por: Zhang, Chaoyun, et al.
Publicado: (2024)
por: Zhang, Chaoyun, et al.
Publicado: (2024)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
por: Huang, Lexiang, et al.
Publicado: (2024)
por: Huang, Lexiang, et al.
Publicado: (2024)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
por: Jain, Kunal, et al.
Publicado: (2024)
por: Jain, Kunal, et al.
Publicado: (2024)
STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
por: Chen, Yinfang, et al.
Publicado: (2025)
por: Chen, Yinfang, et al.
Publicado: (2025)
Towards Cloud Efficiency with Large-scale Workload Characterization
por: Parayil, Anjaly, et al.
Publicado: (2024)
por: Parayil, Anjaly, et al.
Publicado: (2024)
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
por: Dong, Hang, et al.
Publicado: (2024)
por: Dong, Hang, et al.
Publicado: (2024)
Why does Prediction Accuracy Decrease over Time? Uncertain Positive Learning for Cloud Failure Prediction
por: Li, Haozhe, et al.
Publicado: (2024)
por: Li, Haozhe, et al.
Publicado: (2024)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
por: Chu, Xiaoyu, et al.
Publicado: (2025)
por: Chu, Xiaoyu, et al.
Publicado: (2025)
The Vision of Autonomic Computing: Can LLMs Make It a Reality?
por: Zhang, Zhiyang, et al.
Publicado: (2024)
por: Zhang, Zhiyang, et al.
Publicado: (2024)
Huawei Cloud Model-as-a-Service on the CloudMatrix384 SuperPod
por: Xiao, Ao, et al.
Publicado: (2025)
por: Xiao, Ao, et al.
Publicado: (2025)
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
por: Huang, Shaoyuan, et al.
Publicado: (2024)
por: Huang, Shaoyuan, et al.
Publicado: (2024)
XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing
por: Hoefler, Torsten, et al.
Publicado: (2024)
por: Hoefler, Torsten, et al.
Publicado: (2024)
Hybrid-RACA: Hybrid Retrieval-Augmented Composition Assistance for Real-time Text Prediction
por: Xia, Menglin, et al.
Publicado: (2023)
por: Xia, Menglin, et al.
Publicado: (2023)
KCES: A Workflow Containerization Scheduling Scheme Under Cloud-Edge Collaboration Framework
por: Shan, Chenggang, et al.
Publicado: (2024)
por: Shan, Chenggang, et al.
Publicado: (2024)
A Decentralized Microservice Scheduling Approach Using Service Mesh in Cloud-Edge Systems
por: Wen, Yangyang, et al.
Publicado: (2025)
por: Wen, Yangyang, et al.
Publicado: (2025)
Sutradhara: An Intelligent Orchestrator-Engine Co-design for Tool-based Agentic Inference
por: Biswas, Anish, et al.
Publicado: (2026)
por: Biswas, Anish, et al.
Publicado: (2026)
FuxiShuffle: An Adaptive and Resilient Shuffle Service for Distributed Data Processing on Alibaba Cloud
por: Lin, Yuhao, et al.
Publicado: (2026)
por: Lin, Yuhao, et al.
Publicado: (2026)
An Artificial Intelligence Framework for Joint Structural-Temporal Load Forecasting in Cloud Native Platforms
por: Zhang, Qingyuan
Publicado: (2026)
por: Zhang, Qingyuan
Publicado: (2026)
A Scenario-Oriented Benchmark for Assessing AIOps Algorithms in Microservice Management
por: Sun, Yongqian, et al.
Publicado: (2024)
por: Sun, Yongqian, et al.
Publicado: (2024)
Quantifying Autoscaler Vulnerabilities: An Empirical Study of Resource Misallocation Induced by Cloud Infrastructure Faults
por: Park, Gijun
Publicado: (2026)
por: Park, Gijun
Publicado: (2026)
Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
por: Legler, Julian, et al.
Publicado: (2025)
por: Legler, Julian, et al.
Publicado: (2025)
Spider: A BFT Architecture for Geo-Replicated Cloud Services
por: Eischer, Michael, et al.
Publicado: (2024)
por: Eischer, Michael, et al.
Publicado: (2024)
FAILS: A Framework for Automated Collection and Analysis of LLM Service Incidents
por: Battaglini-Fischer, Sándor, et al.
Publicado: (2025)
por: Battaglini-Fischer, Sándor, et al.
Publicado: (2025)
CausalMesh: A Formally Verified Causally Consistent Distributed Cache with Support for Client Migration
por: Zhang, Haoran, et al.
Publicado: (2025)
por: Zhang, Haoran, et al.
Publicado: (2025)
Artifact for Service-Level Energy Modeling and Experimentation for Cloud-Native Microservices
por: Legler, Julian
Publicado: (2026)
por: Legler, Julian
Publicado: (2026)
Tutorial: Object as a Service (OaaS) Serverless Cloud Computing Paradigm
por: Lertpongrujikorn, Pawissanutt, et al.
Publicado: (2024)
por: Lertpongrujikorn, Pawissanutt, et al.
Publicado: (2024)
Benchmarking Different Application Types across Heterogeneous Cloud Compute Services
por: Duggi, Nivedhitha, et al.
Publicado: (2025)
por: Duggi, Nivedhitha, et al.
Publicado: (2025)
Duet instrumentation: An Agentic Approach to Improving Sensitivity in Cloud Service Benchmarking
por: Koch, Sebastian, et al.
Publicado: (2026)
por: Koch, Sebastian, et al.
Publicado: (2026)
FunLess: Functions-as-a-Service for Private Edge Cloud Systems
por: De Palma, Giuseppe, et al.
Publicado: (2024)
por: De Palma, Giuseppe, et al.
Publicado: (2024)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2025)
por: Yang, Zheming, et al.
Publicado: (2025)
Intent-based System Design and Operation
por: Anand, Vaastav, et al.
Publicado: (2025)
por: Anand, Vaastav, et al.
Publicado: (2025)
AI-Driven Multi-Region Provisioning for Cloud Services Using Spot Fleets
por: Fabra, Javier, et al.
Publicado: (2026)
por: Fabra, Javier, et al.
Publicado: (2026)
Minos: Exploiting Cloud Performance Variation with Function-as-a-Service Instance Selection
por: Schirmer, Trever, et al.
Publicado: (2025)
por: Schirmer, Trever, et al.
Publicado: (2025)
Ejemplares similares
-
AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds
por: Chen, Yinfang, et al.
Publicado: (2025) -
Dependency Aware Incident Linking in Large Cloud Systems
por: Ghosh, Supriyo, et al.
Publicado: (2024) -
UFO3: Weaving the Digital Agent Galaxy
por: Zhang, Chaoyun, et al.
Publicado: (2025) -
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
por: Liao, Mengqi, et al.
Publicado: (2026) -
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
por: Jaiswal, Shashwat, et al.
Publicado: (2025)