Taming the Memory Beast: Strategies for Reliable ML Training on Kubernetes
Fuente:
arXiv
Salvato in:
| Autore principale: | Ray, Jaideep |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
di: Dai, Yinwei, et al.
Pubblicazione: (2023)
di: Dai, Yinwei, et al.
Pubblicazione: (2023)
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
di: Li, Shigang, et al.
Pubblicazione: (2019)
di: Li, Shigang, et al.
Pubblicazione: (2019)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
di: Li, Siyuan, et al.
Pubblicazione: (2024)
di: Li, Siyuan, et al.
Pubblicazione: (2024)
Mitigating Temporal Blindness in Kubernetes Autoscaling: An Attention-Double-LSTM Framework
di: Shaikh, Faraz, et al.
Pubblicazione: (2026)
di: Shaikh, Faraz, et al.
Pubblicazione: (2026)
Design and Implementation of an Automated Disaster-recovery System for a Kubernetes Cluster Using LSTM
di: Kim, Ji-Beom, et al.
Pubblicazione: (2024)
di: Kim, Ji-Beom, et al.
Pubblicazione: (2024)
Ilargi: a GPU Compatible Factorized ML Model Training Framework
di: Sun, Wenbo, et al.
Pubblicazione: (2025)
di: Sun, Wenbo, et al.
Pubblicazione: (2025)
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
di: Deng, Yangtao, et al.
Pubblicazione: (2025)
di: Deng, Yangtao, et al.
Pubblicazione: (2025)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
di: Agrawal, Amey, et al.
Pubblicazione: (2024)
PaSE: Parallelization Strategies for Efficient DNN Training
di: Elango, Venmugil
Pubblicazione: (2024)
di: Elango, Venmugil
Pubblicazione: (2024)
Breaking the Memory Wall for Heterogeneous Federated Learning via Progressive Training
di: Wu, Yebo, et al.
Pubblicazione: (2024)
di: Wu, Yebo, et al.
Pubblicazione: (2024)
HeteroSwitch: Characterizing and Taming System-Induced Data Heterogeneity in Federated Learning
di: Kim, Gyudong, et al.
Pubblicazione: (2024)
di: Kim, Gyudong, et al.
Pubblicazione: (2024)
Reducing Memory Contention and I/O Congestion for Disk-based GNN Training
di: Jiang, Qisheng, et al.
Pubblicazione: (2024)
di: Jiang, Qisheng, et al.
Pubblicazione: (2024)
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
di: Waleffe, Roger, et al.
Pubblicazione: (2025)
di: Waleffe, Roger, et al.
Pubblicazione: (2025)
AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
di: Hatanpää, Väinö, et al.
Pubblicazione: (2025)
di: Hatanpää, Väinö, et al.
Pubblicazione: (2025)
Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
di: Fujii, Kazuki, et al.
Pubblicazione: (2024)
di: Fujii, Kazuki, et al.
Pubblicazione: (2024)
Chameleon: Taming Dynamic Operator Sequences for Memory-Intensive LLM Training
di: Wang, Zibo, et al.
Pubblicazione: (2025)
di: Wang, Zibo, et al.
Pubblicazione: (2025)
An Experimental Comparison of Partitioning Strategies for Distributed Graph Neural Network Training
di: Merkel, Nikolai, et al.
Pubblicazione: (2023)
di: Merkel, Nikolai, et al.
Pubblicazione: (2023)
ShardTensor: Domain Parallelism for Scientific Machine Learning
di: Adams, Corey, et al.
Pubblicazione: (2026)
di: Adams, Corey, et al.
Pubblicazione: (2026)
Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
di: Hu, Qinghao, et al.
Pubblicazione: (2025)
di: Hu, Qinghao, et al.
Pubblicazione: (2025)
Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
di: Piao, XinYu, et al.
Pubblicazione: (2021)
di: Piao, XinYu, et al.
Pubblicazione: (2021)
Scaling Deep Learning Research with Kubernetes on the NRP Nautilus HyperCluster
di: Hurt, J. Alex, et al.
Pubblicazione: (2024)
di: Hurt, J. Alex, et al.
Pubblicazione: (2024)
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
di: Li, Haoyang, et al.
Pubblicazione: (2024)
di: Li, Haoyang, et al.
Pubblicazione: (2024)
Evolving HPC services to enable ML workloads on HPE Cray EX
di: Schuppli, Stefano, et al.
Pubblicazione: (2025)
di: Schuppli, Stefano, et al.
Pubblicazione: (2025)
yProv4ML: Effortless Provenance Tracking for Machine Learning Systems
di: Padovani, Gabriele, et al.
Pubblicazione: (2025)
di: Padovani, Gabriele, et al.
Pubblicazione: (2025)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
di: Barrak, Amine, et al.
Pubblicazione: (2025)
di: Barrak, Amine, et al.
Pubblicazione: (2025)
Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
di: Pan, Yi, et al.
Pubblicazione: (2025)
di: Pan, Yi, et al.
Pubblicazione: (2025)
tf.data service: A Case for Disaggregating ML Input Data Processing
di: Audibert, Andrew, et al.
Pubblicazione: (2022)
di: Audibert, Andrew, et al.
Pubblicazione: (2022)
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
di: Ranjan, Aditya K., et al.
Pubblicazione: (2025)
di: Ranjan, Aditya K., et al.
Pubblicazione: (2025)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
di: Zhang, Zongshun, et al.
Pubblicazione: (2025)
di: Zhang, Zongshun, et al.
Pubblicazione: (2025)
SAIR: Cost-Efficient Multi-Stage ML Pipeline Autoscaling via In-Context Reinforcement Learning
di: Su, Jianchang, et al.
Pubblicazione: (2026)
di: Su, Jianchang, et al.
Pubblicazione: (2026)
Improving Automatic Parallel Training via Balanced Memory Workload Optimization
di: Wang, Yujie, et al.
Pubblicazione: (2023)
di: Wang, Yujie, et al.
Pubblicazione: (2023)
Taming the Beast of User-Programmed Transactions on Blockchains: A Declarative Transaction Approach
di: Korchiev, Nodirbek, et al.
Pubblicazione: (2024)
di: Korchiev, Nodirbek, et al.
Pubblicazione: (2024)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
di: Yu, Hanfei, et al.
Pubblicazione: (2025)
di: Yu, Hanfei, et al.
Pubblicazione: (2025)
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
di: Kokolis, Apostolos, et al.
Pubblicazione: (2024)
di: Kokolis, Apostolos, et al.
Pubblicazione: (2024)
A Bring-Your-Own-Model Approach for ML-Driven Storage Placement in Warehouse-Scale Computers
di: Yang, Chenxi, et al.
Pubblicazione: (2025)
di: Yang, Chenxi, et al.
Pubblicazione: (2025)
AGMARL-DKS: An Adaptive Graph-Enhanced Multi-Agent Reinforcement Learning for Dynamic Kubernetes Scheduling
di: Hamzeh, Hamed
Pubblicazione: (2026)
di: Hamzeh, Hamed
Pubblicazione: (2026)
TrainMover: An Interruption-Resilient Runtime for ML Training
di: Lao, ChonLam, et al.
Pubblicazione: (2024)
di: Lao, ChonLam, et al.
Pubblicazione: (2024)
CubicML: Automated ML for Large ML Systems Co-design with ML Prediction of Performance
di: Wen, Wei, et al.
Pubblicazione: (2024)
di: Wen, Wei, et al.
Pubblicazione: (2024)
Enhancing Kubernetes Automated Scheduling with Deep Learning and Reinforcement Techniques for Large-Scale Cloud Computing Optimization
di: Xu, Zheng, et al.
Pubblicazione: (2024)
di: Xu, Zheng, et al.
Pubblicazione: (2024)
Pie: Pooling CPU Memory for LLM Inference
di: Xu, Yi, et al.
Pubblicazione: (2024)
di: Xu, Yi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
di: Dai, Yinwei, et al.
Pubblicazione: (2023) -
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
di: Li, Shigang, et al.
Pubblicazione: (2019) -
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
di: Li, Siyuan, et al.
Pubblicazione: (2024) -
Mitigating Temporal Blindness in Kubernetes Autoscaling: An Attention-Double-LSTM Framework
di: Shaikh, Faraz, et al.
Pubblicazione: (2026) -
Design and Implementation of an Automated Disaster-recovery System for a Kubernetes Cluster Using LSTM
di: Kim, Ji-Beom, et al.
Pubblicazione: (2024)