DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yuanqing, Zhang, Yuchen, Lin, Hao, Hu, Junhao, Zhu, Chunyang, Zhang, Quanlu, Li, Boxun, Dai, Guohao, Yang, Zhi, Cheng, Daning, Zhang, Yunquan, Wang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Asynch-SGBDT: Asynchronous Parallel Stochastic Gradient Boosting Decision Tree based on Parameters Server
von: Daning, Cheng, et al.
Veröffentlicht: (2018)
von: Daning, Cheng, et al.
Veröffentlicht: (2018)
FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
von: Zhu, Zhuoran, et al.
Veröffentlicht: (2025)
von: Zhu, Zhuoran, et al.
Veröffentlicht: (2025)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
von: Kang, Xueze, et al.
Veröffentlicht: (2025)
von: Kang, Xueze, et al.
Veröffentlicht: (2025)
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
von: Huang, Zixiao, et al.
Veröffentlicht: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models
von: Xu, Si, et al.
Veröffentlicht: (2024)
von: Xu, Si, et al.
Veröffentlicht: (2024)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
von: Zhang, Han, et al.
Veröffentlicht: (2026)
von: Zhang, Han, et al.
Veröffentlicht: (2026)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling
von: Pan, Yi, et al.
Veröffentlicht: (2026)
von: Pan, Yi, et al.
Veröffentlicht: (2026)
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
ResiHP: Taming LLM Training Failures with Dynamic Hybrid Parallelism
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
von: Ma, Tenghui, et al.
Veröffentlicht: (2026)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
von: Liang, Antian, et al.
Veröffentlicht: (2025)
von: Liang, Antian, et al.
Veröffentlicht: (2025)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
von: Gu, Diandian, et al.
Veröffentlicht: (2024)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
von: Lu, Yishun, et al.
Veröffentlicht: (2026)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
LiveR: Fine-Grained Elasticity via Live Reconfiguration for Model Training
von: Liu, Haoyuan, et al.
Veröffentlicht: (2026)
von: Liu, Haoyuan, et al.
Veröffentlicht: (2026)
DawnPiper: A Memory-scablable Pipeline Parallel Training Framework
von: Peng, Xuan, et al.
Veröffentlicht: (2025)
von: Peng, Xuan, et al.
Veröffentlicht: (2025)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
von: Ai, Xin, et al.
Veröffentlicht: (2024)
von: Ai, Xin, et al.
Veröffentlicht: (2024)
Accelerating Compound LLM Training Workloads with Maestro
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
von: Hu, Cunchen, et al.
Veröffentlicht: (2024)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training
von: Zheng, Size, et al.
Veröffentlicht: (2026)
von: Zheng, Size, et al.
Veröffentlicht: (2026)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
TASP: Topology-aware Sequence Parallelism
von: Wang, Yida, et al.
Veröffentlicht: (2025)
von: Wang, Yida, et al.
Veröffentlicht: (2025)
CFP: Efficient Optimization of Intra-Operator Parallelism Plans for Large Model Training
von: Hu, Weifang, et al.
Veröffentlicht: (2025)
von: Hu, Weifang, et al.
Veröffentlicht: (2025)
Beyond A Single AI Cluster: A Survey of Decentralized LLM Training
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
von: Dong, Haotian, et al.
Veröffentlicht: (2025)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
von: Wang, Zhixin, et al.
Veröffentlicht: (2025)
von: Wang, Zhixin, et al.
Veröffentlicht: (2025)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
von: Wu, Yongji, et al.
Veröffentlicht: (2024)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
von: Xu, Jiale, et al.
Veröffentlicht: (2025)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2024)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
von: Chen, Chang, et al.
Veröffentlicht: (2025)
von: Chen, Chang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Asynch-SGBDT: Asynchronous Parallel Stochastic Gradient Boosting Decision Tree based on Parameters Server
von: Daning, Cheng, et al.
Veröffentlicht: (2018) -
FUSCO: High-Performance Distributed Data Shuffling via Transformation-Communication Fusion
von: Zhu, Zhuoran, et al.
Veröffentlicht: (2025) -
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
von: Kang, Xueze, et al.
Veröffentlicht: (2025) -
STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
von: Huang, Zixiao, et al.
Veröffentlicht: (2025) -
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)