DHP: Efficient Scaling of MLLM Training with Dynamic Hybrid Parallelism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Niu, Yifan, Xiao, Han, Liu, Dongyi, Zhou, Wei, Li, Jia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
von: Xue, Chunyu, et al.
Veröffentlicht: (2024)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
von: Hagemann, Johannes, et al.
Veröffentlicht: (2023)
von: Hagemann, Johannes, et al.
Veröffentlicht: (2023)
DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
von: Jiang, Chenyu, et al.
Veröffentlicht: (2025)
von: Jiang, Chenyu, et al.
Veröffentlicht: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
von: Bai, Huawei, et al.
Veröffentlicht: (2025)
von: Bai, Huawei, et al.
Veröffentlicht: (2025)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
PaSE: Parallelization Strategies for Efficient DNN Training
von: Elango, Venmugil
Veröffentlicht: (2024)
von: Elango, Venmugil
Veröffentlicht: (2024)
Two-dimensional Sparse Parallelism for Large Scale Deep Learning Recommendation Model Training
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism
von: Polisetty, Sandeep, et al.
Veröffentlicht: (2023)
von: Polisetty, Sandeep, et al.
Veröffentlicht: (2023)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
von: Chen, Ling, et al.
Veröffentlicht: (2026)
von: Chen, Ling, et al.
Veröffentlicht: (2026)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
von: Dash, Sajal, et al.
Veröffentlicht: (2026)
von: Dash, Sajal, et al.
Veröffentlicht: (2026)
FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism
von: Wang, Yujie, et al.
Veröffentlicht: (2024)
von: Wang, Yujie, et al.
Veröffentlicht: (2024)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
von: Jin, Chao, et al.
Veröffentlicht: (2025)
von: Jin, Chao, et al.
Veröffentlicht: (2025)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training
von: Wang, Yuanqing, et al.
Veröffentlicht: (2026)
von: Wang, Yuanqing, et al.
Veröffentlicht: (2026)
Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
von: Li, Haoyang, et al.
Veröffentlicht: (2024)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
Heterogeneous Parallelism for Multimodal Large Language Model Training
von: Karnati, Yashaswi, et al.
Veröffentlicht: (2026)
von: Karnati, Yashaswi, et al.
Veröffentlicht: (2026)
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
von: Bai, Xu, et al.
Veröffentlicht: (2026)
von: Bai, Xu, et al.
Veröffentlicht: (2026)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
ReInc: Scaling Training of Dynamic Graph Neural Networks
von: Guan, Mingyu, et al.
Veröffentlicht: (2025)
von: Guan, Mingyu, et al.
Veröffentlicht: (2025)
Echo: Simulating Distributed Training At Scale
von: Feng, Yicheng, et al.
Veröffentlicht: (2024)
von: Feng, Yicheng, et al.
Veröffentlicht: (2024)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
von: Liu, Xinyi, et al.
Veröffentlicht: (2026)
Unicron: Economizing Self-Healing LLM Training at Scale
von: He, Tao, et al.
Veröffentlicht: (2023)
von: He, Tao, et al.
Veröffentlicht: (2023)
DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
von: Zhao, Xuanlei, et al.
Veröffentlicht: (2024)
Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
von: Zhao, Lingxiao, et al.
Veröffentlicht: (2025)
von: Zhao, Lingxiao, et al.
Veröffentlicht: (2025)
ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
von: Liu, Zedong, et al.
Veröffentlicht: (2025)
von: Liu, Zedong, et al.
Veröffentlicht: (2025)
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
von: Waleffe, Roger, et al.
Veröffentlicht: (2025)
von: Waleffe, Roger, et al.
Veröffentlicht: (2025)
Efficient Distributed MLLM Training with Cornstarch
von: Jang, Insu, et al.
Veröffentlicht: (2025)
von: Jang, Insu, et al.
Veröffentlicht: (2025)
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
von: Wang, Weixun, et al.
Veröffentlicht: (2025)
von: Wang, Weixun, et al.
Veröffentlicht: (2025)
Scaling State-Space Models on Multiple GPUs with Tensor Parallelism
von: Dutt, Anurag, et al.
Veröffentlicht: (2026)
von: Dutt, Anurag, et al.
Veröffentlicht: (2026)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
Reinforcement Learning-Based Dynamic Management of Structured Parallel Farm Skeletons on Serverless Platforms
von: Li, Lanpei, et al.
Veröffentlicht: (2026)
von: Li, Lanpei, et al.
Veröffentlicht: (2026)
Scaling Deep Learning Training with MPMD Pipeline Parallelism
von: Xhebraj, Anxhelo, et al.
Veröffentlicht: (2024)
von: Xhebraj, Anxhelo, et al.
Veröffentlicht: (2024)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
von: Luo, Shuqing, et al.
Veröffentlicht: (2025)
HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
von: Yang, Weihao, et al.
Veröffentlicht: (2025)
von: Yang, Weihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
von: Xue, Chunyu, et al.
Veröffentlicht: (2024) -
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
von: Liu, Dennis, et al.
Veröffentlicht: (2025) -
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
von: Hagemann, Johannes, et al.
Veröffentlicht: (2023) -
DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
von: Jiang, Chenyu, et al.
Veröffentlicht: (2025) -
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)