TrainMover: An Interruption-Resilient Runtime for ML Training
Fuente:
arXiv
Saved in:
| Main Authors: | Lao, ChonLam, Gao, Jiaqi, Cao, Jiamin, Zhang, Zhipeng, Zhang, Pengcheng, Duan, Jiangfei, Zheng, Zhilong, Guan, Yu, Xu, Yichi, Li, Yong, Qian, Zhengping, Akella, Aditya, Yu, Minlan, Zhai, Ennan, Cai, Dennis, Zhou, Jingren |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
by: Lao, ChonLam, et al.
Published: (2024)
by: Lao, ChonLam, et al.
Published: (2024)
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
by: Xi, Shaoke, et al.
Published: (2026)
by: Xi, Shaoke, et al.
Published: (2026)
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning
by: Yoo, Jinsun, et al.
Published: (2025)
by: Yoo, Jinsun, et al.
Published: (2025)
THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
by: Li, Minghao, et al.
Published: (2023)
by: Li, Minghao, et al.
Published: (2023)
Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inference
by: Yi, Ke, et al.
Published: (2024)
by: Yi, Ke, et al.
Published: (2024)
On the Fundamental Limitations of Decentralized Learnable Reward Shaping in Cooperative Multi-Agent Reinforcement Learning
by: Akella, Aditya
Published: (2025)
by: Akella, Aditya
Published: (2025)
Cora: Accelerating Stateful Network Applications with SmartNICs
by: Xi, Shaoke, et al.
Published: (2024)
by: Xi, Shaoke, et al.
Published: (2024)
EROICA: Online Performance Troubleshooting for Large-scale Model Training
by: Guan, Yu, et al.
Published: (2025)
by: Guan, Yu, et al.
Published: (2025)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
by: Liu, Ruitao, et al.
Published: (2026)
by: Liu, Ruitao, et al.
Published: (2026)
Depth-guided NeRF Training via Earth Mover's Distance
by: Rau, Anita, et al.
Published: (2024)
by: Rau, Anita, et al.
Published: (2024)
Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
by: Feng, Weiqi, et al.
Published: (2024)
by: Feng, Weiqi, et al.
Published: (2024)
Comparative Analysis of Time Series Foundation Models for Demographic Forecasting: Enhancing Predictive Accuracy in US Population Dynamics
by: Akella, Aditya, et al.
Published: (2025)
by: Akella, Aditya, et al.
Published: (2025)
Resilience‐Inclusive Capacity Assessment for Railway Networks Under Virtual Coupling via Simulation
by: Jiamin Zhang
Published: (2025)
by: Jiamin Zhang
Published: (2025)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
Introducing Adaptive Continuous Adversarial Training (ACAT) to Enhance ML Robustness
by: elShehaby, Mohamed, et al.
Published: (2024)
by: elShehaby, Mohamed, et al.
Published: (2024)
AutoScout: Structured Optimization for Automating ML System Configuration
by: Shong, Jimmy, et al.
Published: (2026)
by: Shong, Jimmy, et al.
Published: (2026)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
by: Lu, Yishun, et al.
Published: (2026)
by: Lu, Yishun, et al.
Published: (2026)
Accelerating Compound LLM Training Workloads with Maestro
by: Yuan, Xiulong, et al.
Published: (2026)
by: Yuan, Xiulong, et al.
Published: (2026)
Accelerating Learned Image Compression Through Modeling Neural Training Dynamics
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Structure-Guided Adversarial Training of Diffusion Models
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
Resilient Microgrid Formation Considering Communication Interruptions
by: Zhong, Jian, et al.
Published: (2024)
by: Zhong, Jian, et al.
Published: (2024)
DBLP: Phase-Aware Bounded-Loss Transport for Burst-Resilient Distributed ML Training
by: Ma, Zechen, et al.
Published: (2026)
by: Ma, Zechen, et al.
Published: (2026)
The design and hydraulic performance studies of a Tesla valve‐type drip irrigation emitter
by: Tianyu Xu, et al.
Published: (2024)
by: Tianyu Xu, et al.
Published: (2024)
SpinML: Customized Synthetic Data Generation for Private Training of Specialized ML Models
by: Zhang, Jiang, et al.
Published: (2025)
by: Zhang, Jiang, et al.
Published: (2025)
Unicron: Economizing Self-Healing LLM Training at Scale
by: He, Tao, et al.
Published: (2023)
by: He, Tao, et al.
Published: (2023)
Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration
by: Shen, Jucheng, et al.
Published: (2025)
by: Shen, Jucheng, et al.
Published: (2025)
Cyber-Zero: Training Cybersecurity Agents without Runtime
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
Fast Inference for Augmented Large Language Models
by: Shahout, Rana, et al.
Published: (2024)
by: Shahout, Rana, et al.
Published: (2024)
Movers & shakers
Published: (2026)
Published: (2026)
Movers & shakers
Published: (2025)
Published: (2025)
Movers & shakers
Published: (2026)
Published: (2026)
Movers & shakers
Published: (2024)
Published: (2024)
Similar Items
-
EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge
by: Lao, ChonLam, et al.
Published: (2024) -
A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM
by: Xi, Shaoke, et al.
Published: (2026) -
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning
by: Yoo, Jinsun, et al.
Published: (2025) -
THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
by: Li, Minghao, et al.
Published: (2023) -
Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inference
by: Yi, Ke, et al.
Published: (2024)