High Significant Fault Detection in Azure Core Workload Insights
Fuente:
arXiv
Saved in:
| Main Authors: | Lohia, Pranay, Boue, Laurent, Rangappa, Sharath, Agneeswaran, Vijay |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
by: Xie, Zhiqiang, et al.
Published: (2024)
by: Xie, Zhiqiang, et al.
Published: (2024)
AIMeter: Measuring, Analyzing, and Visualizing Energy and Carbon Footprint of AI Workloads
by: Huang, Hongzhen, et al.
Published: (2025)
by: Huang, Hongzhen, et al.
Published: (2025)
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
by: Dong, Hang, et al.
Published: (2024)
by: Dong, Hang, et al.
Published: (2024)
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
by: Zuo, Jingwei, et al.
Published: (2026)
by: Zuo, Jingwei, et al.
Published: (2026)
Hybrid Learning and Optimization-Based Dynamic Scheduling for DL Workloads on Heterogeneous GPU Clusters
by: Dongare, Shruti, et al.
Published: (2025)
by: Dongare, Shruti, et al.
Published: (2025)
Federated Learning with Workload Reduction through Partial Training of Client Models and Entropy-Based Data Selection
by: Shi, Hongrui, et al.
Published: (2024)
by: Shi, Hongrui, et al.
Published: (2024)
Game-Theoretic Deep Reinforcement Learning to Minimize Carbon Emissions and Energy Costs for AI Inference Workloads in Geo-Distributed Data Centers
by: Hogade, Ninad, et al.
Published: (2024)
by: Hogade, Ninad, et al.
Published: (2024)
Role-Based Fault Tolerance System for LLM RL Post-Training
by: Chen, Zhenqian, et al.
Published: (2025)
by: Chen, Zhenqian, et al.
Published: (2025)
FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention
by: Dai, Huangliang, et al.
Published: (2025)
by: Dai, Huangliang, et al.
Published: (2025)
EPSILON: Adaptive Fault Mitigation in Approximate Deep Neural Network using Statistical Signatures
by: Khalil, Khurram, et al.
Published: (2025)
by: Khalil, Khurram, et al.
Published: (2025)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
by: Li, Xiangchen, et al.
Published: (2025)
by: Li, Xiangchen, et al.
Published: (2025)
Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training
by: Liu, Guanliang, et al.
Published: (2026)
by: Liu, Guanliang, et al.
Published: (2026)
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
by: Liu, Ziyue, et al.
Published: (2026)
by: Liu, Ziyue, et al.
Published: (2026)
FedCore: Straggler-Free Federated Learning with Distributed Coresets
by: Guo, Hongpeng, et al.
Published: (2024)
by: Guo, Hongpeng, et al.
Published: (2024)
SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores
by: Mei, Zhiyu, et al.
Published: (2023)
by: Mei, Zhiyu, et al.
Published: (2023)
FedComLoc: Communication-Efficient Distributed Training of Sparse and Quantized Models
by: Yi, Kai, et al.
Published: (2024)
by: Yi, Kai, et al.
Published: (2024)
From Centralized to Decentralized Federated Learning: Theoretical Insights, Privacy Preservation, and Robustness Challenges
by: Li, Qiongxiu, et al.
Published: (2025)
by: Li, Qiongxiu, et al.
Published: (2025)
Redefining Data-Centric Design: A New Approach with a Domain Model and Core Data Ontology for Computational Systems
by: Johnson, William, et al.
Published: (2024)
by: Johnson, William, et al.
Published: (2024)
Cluster Workload Allocation: Semantic Soft Affinity Using Natural Language Processing
by: Sliwko, Leszek, et al.
Published: (2026)
by: Sliwko, Leszek, et al.
Published: (2026)
Cluster Workload Allocation: A Predictive Approach Leveraging Machine Learning Efficiency
by: Sliwko, Leszek
Published: (2025)
by: Sliwko, Leszek
Published: (2025)
Duration-Informed Workload Scheduler
by: Loreti, Daniela, et al.
Published: (2026)
by: Loreti, Daniela, et al.
Published: (2026)
veScale-FSDP: Flexible and High-Performance FSDP at Scale
by: Wang, Zezhou, et al.
Published: (2026)
by: Wang, Zezhou, et al.
Published: (2026)
Incentivizing High-quality Participation From Federated Learning Agents
by: Pang, Jinlong, et al.
Published: (2025)
by: Pang, Jinlong, et al.
Published: (2025)
Scale-up Unlearnable Examples Learning with High-Performance Computing
by: Zhu, Yanfan, et al.
Published: (2025)
by: Zhu, Yanfan, et al.
Published: (2025)
Multi-Tier Labeling and Physics-Informed Learning for Orbital Anomaly Detection at Scale
by: Fu, Yong
Published: (2026)
by: Fu, Yong
Published: (2026)
DeepHYDRA: Resource-Efficient Time-Series Anomaly Detection in Dynamically-Configured Systems
by: Stehle, Franz Kevin, et al.
Published: (2024)
by: Stehle, Franz Kevin, et al.
Published: (2024)
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
by: Hu, Xinyi, et al.
Published: (2026)
by: Hu, Xinyi, et al.
Published: (2026)
FL-GUARD: A Holistic Framework for Run-Time Detection and Recovery of Negative Federated Learning
by: Lin, Hong, et al.
Published: (2024)
by: Lin, Hong, et al.
Published: (2024)
A Feature Engineering Approach for Business Impact-Oriented Failure Detection in Distributed Instant Payment Systems
by: Porcelli, Lorenzo
Published: (2025)
by: Porcelli, Lorenzo
Published: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
by: Cao, Shiyi, et al.
Published: (2024)
by: Cao, Shiyi, et al.
Published: (2024)
Workload Schedulers -- Genesis, Algorithms and Differences
by: Sliwko, Leszek, et al.
Published: (2025)
by: Sliwko, Leszek, et al.
Published: (2025)
Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
by: Dong, Jianbo, et al.
Published: (2024)
by: Dong, Jianbo, et al.
Published: (2024)
Federated Learning for MRI-based BrainAGE: a multicenter study on post-stroke functional outcome prediction
by: Roca, Vincent, et al.
Published: (2025)
by: Roca, Vincent, et al.
Published: (2025)
LiquidGEMM: Hardware-Efficient W4A8 GEMM Kernel for High-Performance LLM Serving
by: Hu, Huanqi, et al.
Published: (2025)
by: Hu, Huanqi, et al.
Published: (2025)
Toward Sustainable GenAI using Generation Directives for Carbon-Friendly Large Language Model Inference
by: Li, Baolin, et al.
Published: (2024)
by: Li, Baolin, et al.
Published: (2024)
Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML
by: John, Chelsea Maria, et al.
Published: (2024)
by: John, Chelsea Maria, et al.
Published: (2024)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
by: Bian, Song, et al.
Published: (2025)
by: Bian, Song, et al.
Published: (2025)
Communication-Efficient Training Workload Balancing for Decentralized Multi-Agent Learning
by: Mohammadabadi, Seyed Mahmoud Sajjadi, et al.
Published: (2024)
by: Mohammadabadi, Seyed Mahmoud Sajjadi, et al.
Published: (2024)
Topology-aware Preemptive Scheduling for Co-located LLM Workloads
by: Zhang, Ping, et al.
Published: (2024)
by: Zhang, Ping, et al.
Published: (2024)
Decentralized Federated Policy Gradient with Byzantine Fault-Tolerance and Provably Fast Convergence
by: Jordan, Philip, et al.
Published: (2024)
by: Jordan, Philip, et al.
Published: (2024)
Similar Items
-
Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
by: Xie, Zhiqiang, et al.
Published: (2024) -
AIMeter: Measuring, Analyzing, and Visualizing Energy and Carbon Footprint of AI Workloads
by: Huang, Hongzhen, et al.
Published: (2025) -
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
by: Dong, Hang, et al.
Published: (2024) -
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
by: Zuo, Jingwei, et al.
Published: (2026) -
Hybrid Learning and Optimization-Based Dynamic Scheduling for DL Workloads on Heterogeneous GPU Clusters
by: Dongare, Shruti, et al.
Published: (2025)