Design and Implementation of an Automated Disaster-recovery System for a Kubernetes Cluster Using LSTM
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Ji-Beom, Choi, Je-Bum, Jung, Eun-Sung |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mitigating Temporal Blindness in Kubernetes Autoscaling: An Attention-Double-LSTM Framework
por: Shaikh, Faraz, et al.
Publicado: (2026)
por: Shaikh, Faraz, et al.
Publicado: (2026)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
por: Li, Siyuan, et al.
Publicado: (2024)
por: Li, Siyuan, et al.
Publicado: (2024)
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
por: Ray, Jaideep
Publicado: (2024)
por: Ray, Jaideep
Publicado: (2024)
Scaling Deep Learning Research with Kubernetes on the NRP Nautilus HyperCluster
por: Hurt, J. Alex, et al.
Publicado: (2024)
por: Hurt, J. Alex, et al.
Publicado: (2024)
KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances
por: Kim, Taeyoon, et al.
Publicado: (2026)
por: Kim, Taeyoon, et al.
Publicado: (2026)
GraNNDis: Efficient Unified Distributed Training Framework for Deep GNNs on Large Clusters
por: Song, Jaeyong, et al.
Publicado: (2023)
por: Song, Jaeyong, et al.
Publicado: (2023)
Pipette: Automatic Fine-grained Large Language Model Training Configurator for Real-World Clusters
por: Yim, Jinkyu, et al.
Publicado: (2024)
por: Yim, Jinkyu, et al.
Publicado: (2024)
QONNECT: A QoS-Aware Orchestration System for Distributed Kubernetes Clusters
por: Aslan, Haci Ismail, et al.
Publicado: (2025)
por: Aslan, Haci Ismail, et al.
Publicado: (2025)
Enhancing Kubernetes Automated Scheduling with Deep Learning and Reinforcement Techniques for Large-Scale Cloud Computing Optimization
por: Xu, Zheng, et al.
Publicado: (2024)
por: Xu, Zheng, et al.
Publicado: (2024)
Weather Prediction Using CNN-LSTM for Time Series Analysis: A Case Study on Delhi Temperature Data
por: Li, Bangyu, et al.
Publicado: (2024)
por: Li, Bangyu, et al.
Publicado: (2024)
Automated Byzantine-Resilient Clustered Decentralized Federated Learning for Battery Intelligence in Connected EVs
por: Bouchiha, Mouhamed Amine, et al.
Publicado: (2026)
por: Bouchiha, Mouhamed Amine, et al.
Publicado: (2026)
Implementation of New Security Features in CMSWEB Kubernetes Cluster at CERN
por: Ali, Aamir, et al.
Publicado: (2024)
por: Ali, Aamir, et al.
Publicado: (2024)
Improving the Efficiency of a Deep Reinforcement Learning-Based Power Management System for HPC Clusters Using Curriculum Learning
por: Budiarjo, Thomas, et al.
Publicado: (2025)
por: Budiarjo, Thomas, et al.
Publicado: (2025)
A Survey on Scheduling Techniques in the Edge Cloud: Issues, Challenges and Future Directions
por: Asghar, Hassan, et al.
Publicado: (2022)
por: Asghar, Hassan, et al.
Publicado: (2022)
Efficient Split Learning LSTM Models for FPGA-based Edge IoT Devices
por: Molina, Romina Soledad, et al.
Publicado: (2025)
por: Molina, Romina Soledad, et al.
Publicado: (2025)
Flexible Clustered Federated Learning for Client-Level Data Distribution Shift
por: Duan, Moming, et al.
Publicado: (2021)
por: Duan, Moming, et al.
Publicado: (2021)
The National Research Platform: Stretched, Multi-Tenant, Scientific Kubernetes Cluster
por: Weitzel, Derek, et al.
Publicado: (2025)
por: Weitzel, Derek, et al.
Publicado: (2025)
Prediction of Brent crude oil price based on LSTM model under the background of low-carbon transition
por: Zhao, Yuwen, et al.
Publicado: (2024)
por: Zhao, Yuwen, et al.
Publicado: (2024)
FedGroup: Efficient Clustered Federated Learning via Decomposed Data-Driven Measure
por: Duan, Moming, et al.
Publicado: (2020)
por: Duan, Moming, et al.
Publicado: (2020)
Asynchronous Federated Clustering with Unknown Number of Clusters
por: Zhang, Yunfan, et al.
Publicado: (2024)
por: Zhang, Yunfan, et al.
Publicado: (2024)
Kubernetes Deployment Options for On-Prem Clusters
por: Bryant, Lincoln, et al.
Publicado: (2024)
por: Bryant, Lincoln, et al.
Publicado: (2024)
Scattered Mixture-of-Experts Implementation
por: Tan, Shawn, et al.
Publicado: (2024)
por: Tan, Shawn, et al.
Publicado: (2024)
Priority Matters: Optimising Kubernetes Clusters Usage with Constraint-Based Pod Packing
por: Christensen, Henrik Daniel, et al.
Publicado: (2025)
por: Christensen, Henrik Daniel, et al.
Publicado: (2025)
Near-Optimal Resilient Aggregation Rules for Distributed Learning Using 1-Center and 1-Mean Clustering with Outliers
por: Yi, Yuhao, et al.
Publicado: (2023)
por: Yi, Yuhao, et al.
Publicado: (2023)
Comprehensive Performance Modeling and System Design Insights for Foundation Models
por: Subramanian, Shashank, et al.
Publicado: (2024)
por: Subramanian, Shashank, et al.
Publicado: (2024)
Communication-Efficient Device Scheduling for Federated Learning Using Lyapunov Optimization
por: Perazzone, Jake B., et al.
Publicado: (2025)
por: Perazzone, Jake B., et al.
Publicado: (2025)
Federated K-means Clustering
por: Garst, Swier, et al.
Publicado: (2023)
por: Garst, Swier, et al.
Publicado: (2023)
Federated Temporal Graph Clustering
por: Zhou, Zihao, et al.
Publicado: (2024)
por: Zhou, Zihao, et al.
Publicado: (2024)
DPMM-CFL: Clustered Federated Learning via Dirichlet Process Mixture Model Nonparametric Clustering
por: Jaramillo-Civill, Mariona, et al.
Publicado: (2025)
por: Jaramillo-Civill, Mariona, et al.
Publicado: (2025)
Federated Automated Feature Engineering
por: Overman, Tom, et al.
Publicado: (2024)
por: Overman, Tom, et al.
Publicado: (2024)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
por: Kim, Geon-Woo, et al.
Publicado: (2025)
por: Kim, Geon-Woo, et al.
Publicado: (2025)
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
por: Kim, Gyudong, et al.
Publicado: (2024)
por: Kim, Gyudong, et al.
Publicado: (2024)
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
por: Lee, Gunjun, et al.
Publicado: (2025)
por: Lee, Gunjun, et al.
Publicado: (2025)
AGMARL-DKS: An Adaptive Graph-Enhanced Multi-Agent Reinforcement Learning for Dynamic Kubernetes Scheduling
por: Hamzeh, Hamed
Publicado: (2026)
por: Hamzeh, Hamed
Publicado: (2026)
Distributed Deep Learning using Stochastic Gradient Staleness
por: Pham, Viet Hoang, et al.
Publicado: (2025)
por: Pham, Viet Hoang, et al.
Publicado: (2025)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
por: Kim, Heehoon, et al.
Publicado: (2026)
por: Kim, Heehoon, et al.
Publicado: (2026)
cuConv: A CUDA Implementation of Convolution for CNN Inference
por: Jordà, Marc, et al.
Publicado: (2021)
por: Jordà, Marc, et al.
Publicado: (2021)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
por: Yan, Ran, et al.
Publicado: (2025)
por: Yan, Ran, et al.
Publicado: (2025)
NL-CPS: Reinforcement Learning-Based Kubernetes Control Plane Placement in Multi-Region Clusters
por: Alam, Sajid, et al.
Publicado: (2026)
por: Alam, Sajid, et al.
Publicado: (2026)
Aryl: An Elastic Cluster Scheduler for Deep Learning
por: Li, Jiamin, et al.
Publicado: (2022)
por: Li, Jiamin, et al.
Publicado: (2022)
Ejemplares similares
-
Mitigating Temporal Blindness in Kubernetes Autoscaling: An Attention-Double-LSTM Framework
por: Shaikh, Faraz, et al.
Publicado: (2026) -
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
por: Li, Siyuan, et al.
Publicado: (2024) -
Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
por: Ray, Jaideep
Publicado: (2024) -
Scaling Deep Learning Research with Kubernetes on the NRP Nautilus HyperCluster
por: Hurt, J. Alex, et al.
Publicado: (2024) -
KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances
por: Kim, Taeyoon, et al.
Publicado: (2026)