Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yang, Zhu, Wenxuan, Quan, Xuehui, Wang, Heyi, Liu, Chang, Wu, Qiyuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Joint Temporal-Structural Representation Learning for Distributed Fault Discrimination in Microservice Architectures
di: Xue, Yihan, et al.
Pubblicazione: (2026)
di: Xue, Yihan, et al.
Pubblicazione: (2026)
Spatiotemporal Traffic Prediction in Distributed Backend Systems via Graph Neural Networks
di: Qiu, Zhimin, et al.
Pubblicazione: (2025)
di: Qiu, Zhimin, et al.
Pubblicazione: (2025)
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
di: Ni, Yinan, et al.
Pubblicazione: (2025)
di: Ni, Yinan, et al.
Pubblicazione: (2025)
A Formal Framework for Predicting Distributed System Performance under Faults (Extended Version)
di: Zhou, Ziwei, et al.
Pubblicazione: (2026)
di: Zhou, Ziwei, et al.
Pubblicazione: (2026)
Approximate Byzantine Fault-Tolerance in Distributed Optimization
di: Liu, Shuo, et al.
Pubblicazione: (2021)
di: Liu, Shuo, et al.
Pubblicazione: (2021)
Asynchronous Fault-Tolerant Language Decidability for Runtime Verification of Distributed Systems
di: Castañeda, Armando, et al.
Pubblicazione: (2025)
di: Castañeda, Armando, et al.
Pubblicazione: (2025)
Decentralized Proactive Model Offloading and Resource Allocation for Split and Federated Learning
di: Huang, Binbin, et al.
Pubblicazione: (2024)
di: Huang, Binbin, et al.
Pubblicazione: (2024)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
di: Zhong, Yuchen, et al.
Pubblicazione: (2024)
di: Zhong, Yuchen, et al.
Pubblicazione: (2024)
Fault-Tolerant Decentralized Distributed Asynchronous Federated Learning with Adaptive Termination Detection
di: Akkinepally, Phani Sahasra, et al.
Pubblicazione: (2025)
di: Akkinepally, Phani Sahasra, et al.
Pubblicazione: (2025)
PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
di: Zhao, Zhixin, et al.
Pubblicazione: (2026)
di: Zhao, Zhixin, et al.
Pubblicazione: (2026)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
di: Liu, Rixin, et al.
Pubblicazione: (2026)
di: Liu, Rixin, et al.
Pubblicazione: (2026)
A Reinforcement Learning-Driven Task Scheduling Algorithm for Multi-Tenant Distributed Systems
di: Zhang, Xiaopei, et al.
Pubblicazione: (2025)
di: Zhang, Xiaopei, et al.
Pubblicazione: (2025)
Asynchronous Fault-Tolerant Distributed Proper Coloring of Graphs
di: Balliu, Alkida, et al.
Pubblicazione: (2024)
di: Balliu, Alkida, et al.
Pubblicazione: (2024)
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
di: Huang, Shaoyuan, et al.
Pubblicazione: (2024)
di: Huang, Shaoyuan, et al.
Pubblicazione: (2024)
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
di: Xiong, Yifan, et al.
Pubblicazione: (2024)
di: Xiong, Yifan, et al.
Pubblicazione: (2024)
Half a Century of Distributed Byzantine Fault-Tolerant Consensus: Design Principles and Evolutionary Pathways
di: Wu, Huanyu, et al.
Pubblicazione: (2024)
di: Wu, Huanyu, et al.
Pubblicazione: (2024)
Embedded Distributed Inference of Deep Neural Networks: A Systematic Review
di: Peccia, Federico Nicolás, et al.
Pubblicazione: (2024)
di: Peccia, Federico Nicolás, et al.
Pubblicazione: (2024)
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection
di: Zhou, Yuhang, et al.
Pubblicazione: (2025)
di: Zhou, Yuhang, et al.
Pubblicazione: (2025)
Domain-Adversarial Transfer Learning for Fault Root Cause Identification in Cloud Computing Systems
di: Fang, Bruce, et al.
Pubblicazione: (2025)
di: Fang, Bruce, et al.
Pubblicazione: (2025)
Maple: A Multi-agent System for Portable Deep Learning across Clusters
di: Wu, Molang, et al.
Pubblicazione: (2025)
di: Wu, Molang, et al.
Pubblicazione: (2025)
Distributed Log-driven Anomaly Detection System based on Evolving Decision Making
di: Tan, Zhuoran, et al.
Pubblicazione: (2025)
di: Tan, Zhuoran, et al.
Pubblicazione: (2025)
Agreement Tasks in Fault-Prone Synchronous Networks of Arbitrary Structure
di: Fraigniaud, Pierre, et al.
Pubblicazione: (2024)
di: Fraigniaud, Pierre, et al.
Pubblicazione: (2024)
Process-Commutative Distributed Objects: From Cryptocurrencies to Byzantine-Fault-Tolerant CRDTs
di: Frey, Davide, et al.
Pubblicazione: (2023)
di: Frey, Davide, et al.
Pubblicazione: (2023)
LA-IMR: Latency-Aware, Predictive In-Memory Routing and Proactive Autoscaling for Tail-Latency-Sensitive Cloud Robotics
di: Seo, Eunil, et al.
Pubblicazione: (2025)
di: Seo, Eunil, et al.
Pubblicazione: (2025)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
di: Xu, Jingwei, et al.
Pubblicazione: (2025)
di: Xu, Jingwei, et al.
Pubblicazione: (2025)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
di: Tanaka, Masahiro, et al.
Pubblicazione: (2025)
di: Tanaka, Masahiro, et al.
Pubblicazione: (2025)
Graph Neural Networks and Reinforcement Learning for Proactive Application Image Placement
di: Makris, Antonios, et al.
Pubblicazione: (2024)
di: Makris, Antonios, et al.
Pubblicazione: (2024)
Proactive and Reactive Autoscaling Techniques for Edge Computing
di: Gupta, Suhrid, et al.
Pubblicazione: (2025)
di: Gupta, Suhrid, et al.
Pubblicazione: (2025)
Taming Cold Starts: Proactive Serverless Scheduling with Model Predictive Control
di: Nguyen, Chanh, et al.
Pubblicazione: (2025)
di: Nguyen, Chanh, et al.
Pubblicazione: (2025)
On Fault Tolerance of Data Storage Systems: A Holistic Perspective
di: Zheng, Mai, et al.
Pubblicazione: (2025)
di: Zheng, Mai, et al.
Pubblicazione: (2025)
BlockRaFT: A Distributed Framework for Fault-Tolerant and Scalable Blockchain Nodes
di: Piduguralla, Manaswini, et al.
Pubblicazione: (2026)
di: Piduguralla, Manaswini, et al.
Pubblicazione: (2026)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
di: Wang, Yisu, et al.
Pubblicazione: (2025)
di: Wang, Yisu, et al.
Pubblicazione: (2025)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
di: Yao, Chenxuan, et al.
Pubblicazione: (2025)
di: Yao, Chenxuan, et al.
Pubblicazione: (2025)
Byzantine Fault-Tolerant Min-Max Optimization
di: Liu, Shuo, et al.
Pubblicazione: (2022)
di: Liu, Shuo, et al.
Pubblicazione: (2022)
Low-Latency Layer-Aware Proactive and Passive Container Migration in Meta Computing
di: Liu, Mengjie, et al.
Pubblicazione: (2024)
di: Liu, Mengjie, et al.
Pubblicazione: (2024)
Hamster: A Fast Synchronous Byzantine Fault Tolerance Protocol
di: Fu, Ximing, et al.
Pubblicazione: (2024)
di: Fu, Ximing, et al.
Pubblicazione: (2024)
MalleTrain: Deep Neural Network Training on Unfillable Supercomputer Nodes
di: Ma, Xiaolong, et al.
Pubblicazione: (2024)
di: Ma, Xiaolong, et al.
Pubblicazione: (2024)
MOSS: A Large-scale Open Microscopic Traffic Simulation System
di: Zhang, Jun, et al.
Pubblicazione: (2024)
di: Zhang, Jun, et al.
Pubblicazione: (2024)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
di: Li, Zixuan, et al.
Pubblicazione: (2026)
di: Li, Zixuan, et al.
Pubblicazione: (2026)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
di: Du, Jiangsu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Joint Temporal-Structural Representation Learning for Distributed Fault Discrimination in Microservice Architectures
di: Xue, Yihan, et al.
Pubblicazione: (2026) -
Spatiotemporal Traffic Prediction in Distributed Backend Systems via Graph Neural Networks
di: Qiu, Zhimin, et al.
Pubblicazione: (2025) -
Predictive-LoRA: A Proactive and Fragmentation-Aware Serverless Inference System for LLMs
di: Ni, Yinan, et al.
Pubblicazione: (2025) -
A Formal Framework for Predicting Distributed System Performance under Faults (Extended Version)
di: Zhou, Ziwei, et al.
Pubblicazione: (2026) -
Approximate Byzantine Fault-Tolerance in Distributed Optimization
di: Liu, Shuo, et al.
Pubblicazione: (2021)