ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gandhi, Swapnil, Zhao, Mark, Skiadopoulos, Athinagoras, Kozyrakis, Christos |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
FailSafe: High-performance Resilient Serving
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
von: Xu, Ziyi, et al.
Veröffentlicht: (2025)
cedar: Optimized and Unified Machine Learning Input Data Pipelines
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
von: Zhao, Mark, et al.
Veröffentlicht: (2024)
Regulating Branch Parallelism in LLM Serving
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
Cloud Atlas: Efficient Fault Localization for Cloud Systems using Language Models and Causal Insight
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2024)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
von: Chen, Ling, et al.
Veröffentlicht: (2026)
von: Chen, Ling, et al.
Veröffentlicht: (2026)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline
von: Sheng, Guangming, et al.
Veröffentlicht: (2024)
von: Sheng, Guangming, et al.
Veröffentlicht: (2024)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2024)
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2024)
SkipPipe: Partial and Reordered Pipelining Framework for Training LLMs in Heterogeneous Networks
von: Blagoev, Nikolay, et al.
Veröffentlicht: (2025)
von: Blagoev, Nikolay, et al.
Veröffentlicht: (2025)
A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training
von: Barley, Daniel, et al.
Veröffentlicht: (2026)
von: Barley, Daniel, et al.
Veröffentlicht: (2026)
ReInc: Scaling Training of Dynamic Graph Neural Networks
von: Guan, Mingyu, et al.
Veröffentlicht: (2025)
von: Guan, Mingyu, et al.
Veröffentlicht: (2025)
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
von: Huang, Mincong, et al.
Veröffentlicht: (2024)
von: Huang, Mincong, et al.
Veröffentlicht: (2024)
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2024)
IPA: Inference Pipeline Adaptation to Achieve High Accuracy and Cost-Efficiency
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
von: Ghafouri, Saeid, et al.
Veröffentlicht: (2023)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
von: Wu, Houming, et al.
Veröffentlicht: (2024)
von: Wu, Houming, et al.
Veröffentlicht: (2024)
DNN-Powered MLOps Pipeline Optimization for Large Language Models: A Framework for Automated Deployment and Resource Management
von: Krishnamoorthy, Mahesh Vaijainthymala, et al.
Veröffentlicht: (2025)
von: Krishnamoorthy, Mahesh Vaijainthymala, et al.
Veröffentlicht: (2025)
Understanding Stragglers in Large Model Training Using What-if Analysis
von: Lin, Jinkun, et al.
Veröffentlicht: (2025)
von: Lin, Jinkun, et al.
Veröffentlicht: (2025)
Strata: Hierarchical Context Caching for Long Context Language Model Serving
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2025)
von: Xie, Zhiqiang, et al.
Veröffentlicht: (2025)
Reducing Energy Bloat in Large Model Training
von: Chung, Jae-Won, et al.
Veröffentlicht: (2023)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2023)
Scaling State-Space Models on Multiple GPUs with Tensor Parallelism
von: Dutt, Anurag, et al.
Veröffentlicht: (2026)
von: Dutt, Anurag, et al.
Veröffentlicht: (2026)
Practical Performance Guarantees for Pipelined DNN Inference
von: Archer, Aaron, et al.
Veröffentlicht: (2023)
von: Archer, Aaron, et al.
Veröffentlicht: (2023)
Governing Cloud Data Pipelines with Agentic AI
von: Kirubakaran, Aswathnarayan Muthukrishnan, et al.
Veröffentlicht: (2025)
von: Kirubakaran, Aswathnarayan Muthukrishnan, et al.
Veröffentlicht: (2025)
Nesterov Method for Asynchronous Pipeline Parallel Optimization
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2025)
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2025)
SAIR: Cost-Efficient Multi-Stage ML Pipeline Autoscaling via In-Context Reinforcement Learning
von: Su, Jianchang, et al.
Veröffentlicht: (2026)
von: Su, Jianchang, et al.
Veröffentlicht: (2026)
Heterogeneous Parallelism for Multimodal Large Language Model Training
von: Karnati, Yashaswi, et al.
Veröffentlicht: (2026)
von: Karnati, Yashaswi, et al.
Veröffentlicht: (2026)
PiPar: Pipeline Parallelism for Collaborative Machine Learning
von: Zhang, Zihan, et al.
Veröffentlicht: (2022)
von: Zhang, Zihan, et al.
Veröffentlicht: (2022)
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations
von: Wang, Ziyao, et al.
Veröffentlicht: (2024)
von: Wang, Ziyao, et al.
Veröffentlicht: (2024)
EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
von: Yu, Zhongzhi, et al.
Veröffentlicht: (2024)
von: Yu, Zhongzhi, et al.
Veröffentlicht: (2024)
Optimizing Large Model Training through Overlapped Activation Recomputation
von: Chen, Ping, et al.
Veröffentlicht: (2024)
von: Chen, Ping, et al.
Veröffentlicht: (2024)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
von: Hagemann, Johannes, et al.
Veröffentlicht: (2023)
von: Hagemann, Johannes, et al.
Veröffentlicht: (2023)
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data
von: Kapadia, Shashank, et al.
Veröffentlicht: (2026)
von: Kapadia, Shashank, et al.
Veröffentlicht: (2026)
AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2026)
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025) -
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024) -
FailSafe: High-performance Resilient Serving
von: Xu, Ziyi, et al.
Veröffentlicht: (2025) -
cedar: Optimized and Unified Machine Learning Input Data Pipelines
von: Zhao, Mark, et al.
Veröffentlicht: (2024) -
Regulating Branch Parallelism in LLM Serving
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)