BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Tao, Chen, Runqing, Wang, Pinghui, Zhao, Junzhou, Liu, Jiongzhou, Han, Shujie, Liu, Yi, Xu, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
by: Darzi, Erfan, et al.
Published: (2025)
by: Darzi, Erfan, et al.
Published: (2025)
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
by: Suman, Shekhar, et al.
Published: (2026)
by: Suman, Shekhar, et al.
Published: (2026)
SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
by: Ji, Shixin, et al.
Published: (2024)
by: Ji, Shixin, et al.
Published: (2024)
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
by: Wilkins, Grant, et al.
Published: (2026)
by: Wilkins, Grant, et al.
Published: (2026)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
by: Xiong, Yifan, et al.
Published: (2024)
by: Xiong, Yifan, et al.
Published: (2024)
Optimization Opportunities for Cloud-Based Data Pipeline Infrastructures
by: Jablonski, Johannes, et al.
Published: (2026)
by: Jablonski, Johannes, et al.
Published: (2026)
Deoxys: A Causal Inference Engine for Unhealthy Node Mitigation in Large-scale Cloud Infrastructure
by: Zhang, Chaoyun, et al.
Published: (2024)
by: Zhang, Chaoyun, et al.
Published: (2024)
Analysis of Server Throughput For Managed Big Data Analytics Frameworks
by: Anagnostakis, Emmanouil, et al.
Published: (2025)
by: Anagnostakis, Emmanouil, et al.
Published: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
by: Chen, Jiabin, et al.
Published: (2024)
by: Chen, Jiabin, et al.
Published: (2024)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
by: Chu, Xiaoyu, et al.
Published: (2025)
by: Chu, Xiaoyu, et al.
Published: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
by: Han, Yunhe, et al.
Published: (2026)
by: Han, Yunhe, et al.
Published: (2026)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
by: Li, Shengwei, et al.
Published: (2023)
by: Li, Shengwei, et al.
Published: (2023)
M$^2$-MFP: A Multi-Scale and Multi-Level Memory Failure Prediction Framework for Reliable Cloud Infrastructure
by: Xie, Hongyi, et al.
Published: (2025)
by: Xie, Hongyi, et al.
Published: (2025)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
by: Sakip, Akhmed, et al.
Published: (2026)
by: Sakip, Akhmed, et al.
Published: (2026)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
A Multi-Server Information-Sharing Environment for Cross-Party Collaboration on A Private Cloud
by: Zhang, Jianping, et al.
Published: (2024)
by: Zhang, Jianping, et al.
Published: (2024)
GPU-Accelerated Batch-Dynamic Subgraph Matching
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
Are Bus-Mounted Edge Servers Feasible?
by: Li, Xuezhi, et al.
Published: (2025)
by: Li, Xuezhi, et al.
Published: (2025)
Quantifying Autoscaler Vulnerabilities: An Empirical Study of Resource Misallocation Induced by Cloud Infrastructure Faults
by: Park, Gijun
Published: (2026)
by: Park, Gijun
Published: (2026)
Batch Denoising for AIGC Service Provisioning in Wireless Edge Networks
by: Xu, Jinghang, et al.
Published: (2025)
by: Xu, Jinghang, et al.
Published: (2025)
A Dynamic Approach to Load Balancing in Cloud Infrastructure: Enhancing Energy Efficiency and Resource Utilization
by: Sakib, Shadman, et al.
Published: (2025)
by: Sakib, Shadman, et al.
Published: (2025)
AgentFlow: Resilient Adaptive Cloud-Edge Framework for Multi-Agent Coordination
by: Chen, Ching Han, et al.
Published: (2025)
by: Chen, Ching Han, et al.
Published: (2025)
Practical Federated Learning without a Server
by: Dhasade, Akash, et al.
Published: (2025)
by: Dhasade, Akash, et al.
Published: (2025)
Bridging Global Frameworks: Governance Strategies Behind Cisco Common Control Framework v4.0 for Scalable Cloud Compliance
by: Sonkar, Nishant
Published: (2025)
by: Sonkar, Nishant
Published: (2025)
Pier: Efficient Large Language Model pretraining with Relaxed Global Communication
by: Fan, Shuyuan, et al.
Published: (2025)
by: Fan, Shuyuan, et al.
Published: (2025)
Pico-Cloud: Cloud Infrastructure for Tiny Edge Devices
by: Guri, Mordechai
Published: (2025)
by: Guri, Mordechai
Published: (2025)
Repurposing of the Run 2 CMS High Level Trigger Infrastructure as a Cloud Resource for Offline Computing
by: Mascheroni, Marco, et al.
Published: (2024)
by: Mascheroni, Marco, et al.
Published: (2024)
Leveraging Public Cloud Infrastructure for Real-time Connected Vehicle Speed Advisory at a Signalized Corridor
by: Deng, Hsien-Wen, et al.
Published: (2024)
by: Deng, Hsien-Wen, et al.
Published: (2024)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025)
by: Zhang, Mingjun, et al.
Published: (2025)
Humas: A Heterogeneity- and Upgrade-aware Microservice Auto-scaling Framework in Large-scale Data Centers
by: Hua, Qin, et al.
Published: (2024)
by: Hua, Qin, et al.
Published: (2024)
The SAP Cloud Infrastructure Dataset: A Reality Check of Scheduling and Placement of VMs in Cloud Computing
by: Uhlig, Arno, et al.
Published: (2025)
by: Uhlig, Arno, et al.
Published: (2025)
Experimental Analysis of Server-Side Caching for Web Performance
by: Umar, Mohammad, et al.
Published: (2026)
by: Umar, Mohammad, et al.
Published: (2026)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
by: Zhong, Yuchen, et al.
Published: (2023)
by: Zhong, Yuchen, et al.
Published: (2023)
ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies
by: Liu, Yuze, et al.
Published: (2026)
by: Liu, Yuze, et al.
Published: (2026)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
by: Tharwani, Jay, et al.
Published: (2025)
by: Tharwani, Jay, et al.
Published: (2025)
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
by: Fan, Yuankai, et al.
Published: (2025)
by: Fan, Yuankai, et al.
Published: (2025)
KCES: A Workflow Containerization Scheduling Scheme Under Cloud-Edge Collaboration Framework
by: Shan, Chenggang, et al.
Published: (2024)
by: Shan, Chenggang, et al.
Published: (2024)
Similar Items
-
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
by: Darzi, Erfan, et al.
Published: (2025) -
Literature Study on Operational Data Analytics Frameworks in Large-scale Computing Infrastructures
by: Suman, Shekhar, et al.
Published: (2026) -
SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
by: Ji, Shixin, et al.
Published: (2024) -
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
by: Wilkins, Grant, et al.
Published: (2026) -
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
by: Liu, Yi, et al.
Published: (2025)