Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Zhiqiang, Zheng, Yujia, Ottens, Lizi, Zhang, Kun, Kozyrakis, Christos, Mace, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FailSafe: High-performance Resilient Serving
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2025)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
von: Xu, Minxian, et al.
Veröffentlicht: (2026)
von: Xu, Minxian, et al.
Veröffentlicht: (2026)
Domain-Adversarial Transfer Learning for Fault Root Cause Identification in Cloud Computing Systems
von: Fang, Bruce, et al.
Veröffentlicht: (2025)
von: Fang, Bruce, et al.
Veröffentlicht: (2025)
ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
Flora: Efficient Cloud Resource Selection for Big Data Processing via Job Classification
von: Will, Jonathan, et al.
Veröffentlicht: (2025)
von: Will, Jonathan, et al.
Veröffentlicht: (2025)
A Self-Healing and Fault-Tolerant Cloud-based Digital Twin Processing Management Model
von: Saxena, Deepika, et al.
Veröffentlicht: (2025)
von: Saxena, Deepika, et al.
Veröffentlicht: (2025)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
Adaptive Fault Tolerance Mechanisms of Large Language Models in Cloud Computing Environments
von: Jin, Yihong, et al.
Veröffentlicht: (2025)
von: Jin, Yihong, et al.
Veröffentlicht: (2025)
Efficient Profit Maximization in Reliability Concerned Static Vehicular Cloud System
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2023)
von: Sarkar, Suvarthi, et al.
Veröffentlicht: (2023)
Huawei Cloud Model-as-a-Service on the CloudMatrix384 SuperPod
von: Xiao, Ao, et al.
Veröffentlicht: (2025)
von: Xiao, Ao, et al.
Veröffentlicht: (2025)
Quantifying Autoscaler Vulnerabilities: An Empirical Study of Resource Misallocation Induced by Cloud Infrastructure Faults
von: Park, Gijun
Veröffentlicht: (2026)
von: Park, Gijun
Veröffentlicht: (2026)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2024)
Cloud-Based AI Systems: Leveraging Large Language Models for Intelligent Fault Detection and Autonomous Self-Healing
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
von: Ji, Cheng, et al.
Veröffentlicht: (2025)
Efficient Probabilistic Workflow Scheduling for IaaS Clouds
von: Russo, Gabriele Russo, et al.
Veröffentlicht: (2024)
von: Russo, Gabriele Russo, et al.
Veröffentlicht: (2024)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
von: Dong, Jiangwen, et al.
Veröffentlicht: (2025)
von: Dong, Jiangwen, et al.
Veröffentlicht: (2025)
Cloud Revolution: Tracing the Origins and Rise of Cloud Computing
von: Gurung, Deepa, et al.
Veröffentlicht: (2025)
von: Gurung, Deepa, et al.
Veröffentlicht: (2025)
LaissezCloud: Continuous Resource Renegotiation for the Public Cloud
von: Harith, Tejas, et al.
Veröffentlicht: (2026)
von: Harith, Tejas, et al.
Veröffentlicht: (2026)
An Efficient Approach for Energy Conservation in Cloud Computing Environment
von: Pande, Sohan Kumar, et al.
Veröffentlicht: (2025)
von: Pande, Sohan Kumar, et al.
Veröffentlicht: (2025)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
von: Chang, Tzu-Tao, et al.
Veröffentlicht: (2025)
Cloud Native System for LLM Inference Serving
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
von: Xu, Minxian, et al.
Veröffentlicht: (2025)
EES-CND: Collaborative Neural Decision-Making for Drift-Aware Fault-Tolerant Edge-Cloud Service Placement
von: Herabad, Mohammadsadeq Garshasbi, et al.
Veröffentlicht: (2026)
von: Herabad, Mohammadsadeq Garshasbi, et al.
Veröffentlicht: (2026)
Development of a Cloud-Based Payroll Management System
von: Aina, Adeyemi, et al.
Veröffentlicht: (2025)
von: Aina, Adeyemi, et al.
Veröffentlicht: (2025)
Towards an Adaptive Runtime System for Cloud-Native HPC
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
von: Bhosale, Aditya, et al.
Veröffentlicht: (2026)
Byzantine Fault Tolerant Causal Ordering
von: Misra, Anshuman, et al.
Veröffentlicht: (2021)
von: Misra, Anshuman, et al.
Veröffentlicht: (2021)
Availability Modeling for Blockchain Provisioning in Private Clouds
von: Dantas, J, et al.
Veröffentlicht: (2025)
von: Dantas, J, et al.
Veröffentlicht: (2025)
H-EYE: Holistic Resource Modeling and Management for Diversely Scaled Edge-Cloud Systems
von: Dagli, Ismet, et al.
Veröffentlicht: (2024)
von: Dagli, Ismet, et al.
Veröffentlicht: (2024)
Generating representative macrobenchmark microservice systems from distributed traces with Palette
von: Anand, Vaastav, et al.
Veröffentlicht: (2025)
von: Anand, Vaastav, et al.
Veröffentlicht: (2025)
Distributed Resource Selection for Self-Organising Cloud-Edge Systems
von: Renau, Quentin, et al.
Veröffentlicht: (2025)
von: Renau, Quentin, et al.
Veröffentlicht: (2025)
Orchestrating Mixed-Criticality Cloud Workloads in Reconfigurable Manufacturing Systems
von: Barletta, Marco, et al.
Veröffentlicht: (2024)
von: Barletta, Marco, et al.
Veröffentlicht: (2024)
DSPE: Profit Maximization in Edge-Cloud Storage System using Dynamic Space Partitioning with Erasure Code
von: Roy, Shubhradeep, et al.
Veröffentlicht: (2025)
von: Roy, Shubhradeep, et al.
Veröffentlicht: (2025)
SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
Failure-Resilient and Carbon-Efficient Deployment of Microservices over the Cloud-Edge Continuum
von: Ponce, Francisco, et al.
Veröffentlicht: (2026)
von: Ponce, Francisco, et al.
Veröffentlicht: (2026)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
von: Yu, Shibo, et al.
Veröffentlicht: (2025)
Managing Forensic Recovery in the Cloud
von: Weir, George R. S., et al.
Veröffentlicht: (2024)
von: Weir, George R. S., et al.
Veröffentlicht: (2024)
Cloud-Enabled Virtual Prototypes
von: Kraus, Tim, et al.
Veröffentlicht: (2025)
von: Kraus, Tim, et al.
Veröffentlicht: (2025)
Cloud abstractions for AI workloads
von: Canini, Marco, et al.
Veröffentlicht: (2025)
von: Canini, Marco, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FailSafe: High-performance Resilient Serving
von: Xu, Ziyi, et al.
Veröffentlicht: (2025) -
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024) -
Strata: Hierarchical Context Caching for Long Context Language Model Serving
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2025) -
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
von: Mathews, Dhanya R, et al.
Veröffentlicht: (2025) -
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
von: Xu, Minxian, et al.
Veröffentlicht: (2026)