Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
Fuente:
arXiv
Saved in:
| Main Authors: | Duan, Jiangfei, Song, Ziang, Miao, Xupeng, Xi, Xiaoli, Lin, Dahua, Xu, Harry, Zhang, Minjia, Jia, Zhihao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
by: Wu, Yongji, et al.
Published: (2025)
by: Wu, Yongji, et al.
Published: (2025)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
by: K., Prashanthi S., et al.
Published: (2025)
by: K., Prashanthi S., et al.
Published: (2025)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
by: Chen, Jiefei, et al.
Published: (2026)
by: Chen, Jiefei, et al.
Published: (2026)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
by: Tayal, Mumuksh, et al.
Published: (2025)
by: Tayal, Mumuksh, et al.
Published: (2025)
Atlas: Hierarchical Partitioning for Quantum Circuit Simulation on GPUs (Extended Version)
by: Xu, Mingkuan, et al.
Published: (2024)
by: Xu, Mingkuan, et al.
Published: (2024)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
by: Jeon, Byungsoo, et al.
Published: (2024)
by: Jeon, Byungsoo, et al.
Published: (2024)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
by: Chen, Chang, et al.
Published: (2025)
by: Chen, Chang, et al.
Published: (2025)
A System for Microserving of LLMs
by: Jin, Hongyi, et al.
Published: (2024)
by: Jin, Hongyi, et al.
Published: (2024)
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
by: Zhang, Shiwei, et al.
Published: (2024)
by: Zhang, Shiwei, et al.
Published: (2024)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
by: Shu, Zhihao, et al.
Published: (2026)
by: Shu, Zhihao, et al.
Published: (2026)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
by: Lian, Xinyu, et al.
Published: (2024)
by: Lian, Xinyu, et al.
Published: (2024)
EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration
by: Du, Jiangsu, et al.
Published: (2025)
by: Du, Jiangsu, et al.
Published: (2025)
Training DNN Models over Heterogeneous Clusters with Optimal Performance
by: Nie, Chengyi, et al.
Published: (2024)
by: Nie, Chengyi, et al.
Published: (2024)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
by: K., Prashanthi S., et al.
Published: (2023)
by: K., Prashanthi S., et al.
Published: (2023)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
by: Zhong, Yuchen, et al.
Published: (2023)
by: Zhong, Yuchen, et al.
Published: (2023)
Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
by: Yu, Enda, et al.
Published: (2024)
by: Yu, Enda, et al.
Published: (2024)
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
by: Jiang, Lijuan, et al.
Published: (2025)
by: Jiang, Lijuan, et al.
Published: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
by: Shi, Xiaoxiang, et al.
Published: (2025)
by: Shi, Xiaoxiang, et al.
Published: (2025)
LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
by: Nie, Xiaonan, et al.
Published: (2024)
by: Nie, Xiaonan, et al.
Published: (2024)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
EaCO: Resource Sharing Dynamics and Its Impact on Energy Efficiency for DNN Training
by: Haghshenas, Kawsar, et al.
Published: (2024)
by: Haghshenas, Kawsar, et al.
Published: (2024)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
by: Svedas, Jonas, et al.
Published: (2025)
by: Svedas, Jonas, et al.
Published: (2025)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
by: Ma, Tenghui, et al.
Published: (2026)
by: Ma, Tenghui, et al.
Published: (2026)
Memory Efficient and Staleness Free Pipeline Parallel DNN Training Framework with Improved Convergence Speed
by: Dutta, Ankita, et al.
Published: (2025)
by: Dutta, Ankita, et al.
Published: (2025)
Infer-EDGE: Dynamic DNN Inference Optimization in 'Just-in-time' Edge-AI Implementations
by: Mounesan, Motahare, et al.
Published: (2025)
by: Mounesan, Motahare, et al.
Published: (2025)
Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
TiMePReSt: Time and Memory Efficient Pipeline Parallel DNN Training with Removed Staleness
by: Dutta, Ankita, et al.
Published: (2024)
by: Dutta, Ankita, et al.
Published: (2024)
PowerTrain: Fast, Generalizable Time and Power Prediction Models to Optimize DNN Training on Accelerated Edges
by: K., Prashanthi S., et al.
Published: (2024)
by: K., Prashanthi S., et al.
Published: (2024)
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
by: Mei, Yixuan, et al.
Published: (2024)
by: Mei, Yixuan, et al.
Published: (2024)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
by: Gupta, Ahan, et al.
Published: (2026)
by: Gupta, Ahan, et al.
Published: (2026)
Improving Automatic Parallel Training via Balanced Memory Workload Optimization
by: Wang, Yujie, et al.
Published: (2023)
by: Wang, Yujie, et al.
Published: (2023)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
by: Lin, Zheng, et al.
Published: (2024)
by: Lin, Zheng, et al.
Published: (2024)
Ding-Dong Ditch: Peeking Into Spot Instance Availability
by: Kim, Kyumin, et al.
Published: (2026)
by: Kim, Kyumin, et al.
Published: (2026)
Low-Latency Layer-Aware Proactive and Passive Container Migration in Meta Computing
by: Liu, Mengjie, et al.
Published: (2024)
by: Liu, Mengjie, et al.
Published: (2024)
Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
PaSE: Parallelization Strategies for Efficient DNN Training
by: Elango, Venmugil
Published: (2024)
by: Elango, Venmugil
Published: (2024)
Similar Items
-
MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
by: Duan, Jiangfei, et al.
Published: (2024) -
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
by: Wu, Yongji, et al.
Published: (2025) -
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
by: K., Prashanthi S., et al.
Published: (2025) -
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
by: Duan, Jiangfei, et al.
Published: (2024) -
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
by: Chen, Jiefei, et al.
Published: (2026)