Scaling Deep Learning Training with MPMD Pipeline Parallelism
Fuente:
arXiv
Saved in:
| Main Authors: | Xhebraj, Anxhelo, Lee, Sean, Chen, Hanfeng, Grover, Vinod |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Morphling: Fast, Fused, and Flexible GNN Training at Scale
by: Anubhab, et al.
Published: (2025)
by: Anubhab, et al.
Published: (2025)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
by: Chen, Ling, et al.
Published: (2026)
by: Chen, Ling, et al.
Published: (2026)
Pipeline Parallelism with Controllable Memory
by: Qi, Penghui, et al.
Published: (2024)
by: Qi, Penghui, et al.
Published: (2024)
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
by: Jin, Hongyi, et al.
Published: (2026)
by: Jin, Hongyi, et al.
Published: (2026)
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
by: Huang, Mincong, et al.
Published: (2024)
by: Huang, Mincong, et al.
Published: (2024)
Two-dimensional Sparse Parallelism for Large Scale Deep Learning Recommendation Model Training
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
by: Liu, Ruitao, et al.
Published: (2026)
by: Liu, Ruitao, et al.
Published: (2026)
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
by: Li, Youjie, et al.
Published: (2025)
by: Li, Youjie, et al.
Published: (2025)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
by: Zhang, Geng, et al.
Published: (2025)
by: Zhang, Geng, et al.
Published: (2025)
Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
by: Jin, Tian, et al.
Published: (2025)
by: Jin, Tian, et al.
Published: (2025)
PiPar: Pipeline Parallelism for Collaborative Machine Learning
by: Zhang, Zihan, et al.
Published: (2022)
by: Zhang, Zihan, et al.
Published: (2022)
Learned Cost Model for Placement on Reconfigurable Dataflow Hardware
by: Guha, Etash, et al.
Published: (2025)
by: Guha, Etash, et al.
Published: (2025)
LOOPer: A Learned Automatic Code Optimizer For Polyhedral Compilers
by: Merouani, Massinissa, et al.
Published: (2024)
by: Merouani, Massinissa, et al.
Published: (2024)
PartIR: Composing SPMD Partitioning Strategies for Machine Learning
by: Alabed, Sami, et al.
Published: (2024)
by: Alabed, Sami, et al.
Published: (2024)
A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training
by: Barley, Daniel, et al.
Published: (2026)
by: Barley, Daniel, et al.
Published: (2026)
Charon: A Unified and Fine-Grained Simulator for Large-Scale LLM Training and Inference
by: Yang, Mengtian, et al.
Published: (2026)
by: Yang, Mengtian, et al.
Published: (2026)
Nesterov Method for Asynchronous Pipeline Parallel Optimization
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
DHP: Efficient Scaling of MLLM Training with Dynamic Hybrid Parallelism
by: Niu, Yifan, et al.
Published: (2026)
by: Niu, Yifan, et al.
Published: (2026)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
by: Hagemann, Johannes, et al.
Published: (2023)
by: Hagemann, Johannes, et al.
Published: (2023)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
by: Cheng, Xinhao, et al.
Published: (2025)
by: Cheng, Xinhao, et al.
Published: (2025)
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
by: Jiang, Zhida, et al.
Published: (2026)
by: Jiang, Zhida, et al.
Published: (2026)
Theoretical Foundations of GPU-Native Compilation for Rapid Code Iteration
by: Metinov, Adilet, et al.
Published: (2025)
by: Metinov, Adilet, et al.
Published: (2025)
Data-efficient Performance Modeling via Pre-training
by: Liu, Chunting, et al.
Published: (2025)
by: Liu, Chunting, et al.
Published: (2025)
VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination
by: Hu, Muyan, et al.
Published: (2026)
by: Hu, Muyan, et al.
Published: (2026)
GPU-Accelerated Synthesis of Mixed-Boolean Arithmetic: Beyond Caching
by: Bathie, Gabriel, et al.
Published: (2026)
by: Bathie, Gabriel, et al.
Published: (2026)
AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
by: Ajanthan, Thalaiyasingam, et al.
Published: (2026)
by: Ajanthan, Thalaiyasingam, et al.
Published: (2026)
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
by: Charles, Zachary, et al.
Published: (2025)
by: Charles, Zachary, et al.
Published: (2025)
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
by: Bai, Xu, et al.
Published: (2026)
by: Bai, Xu, et al.
Published: (2026)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
by: Liu, Dennis, et al.
Published: (2025)
by: Liu, Dennis, et al.
Published: (2025)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
by: Wu, Houming, et al.
Published: (2024)
by: Wu, Houming, et al.
Published: (2024)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2025)
by: Arfeen, Daiyaan, et al.
Published: (2025)
GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism
by: Polisetty, Sandeep, et al.
Published: (2023)
by: Polisetty, Sandeep, et al.
Published: (2023)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
by: Dash, Sajal, et al.
Published: (2026)
by: Dash, Sajal, et al.
Published: (2026)
SENTINEL: Stagewise Integrity Verification for Pipeline Parallel Decentralized Training
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
NPB-Rust: NAS Parallel Benchmarks in Rust
by: Martins, Eduardo M., et al.
Published: (2025)
by: Martins, Eduardo M., et al.
Published: (2025)
OMPILOT: Harnessing Transformer Models for Auto Parallelization to Shared Memory Computing Paradigms
by: Bhattacharjee, Arijit, et al.
Published: (2025)
by: Bhattacharjee, Arijit, et al.
Published: (2025)
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
by: Butler, Branden, et al.
Published: (2024)
by: Butler, Branden, et al.
Published: (2024)
SkipPipe: Partial and Reordered Pipelining Framework for Training LLMs in Heterogeneous Networks
by: Blagoev, Nikolay, et al.
Published: (2025)
by: Blagoev, Nikolay, et al.
Published: (2025)
An MLIR Lowering Pipeline for Stencils at Wafer-Scale
by: Stawinoga, Nicolai, et al.
Published: (2026)
by: Stawinoga, Nicolai, et al.
Published: (2026)
Similar Items
-
Morphling: Fast, Fused, and Flexible GNN Training at Scale
by: Anubhab, et al.
Published: (2025) -
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
by: Chen, Ling, et al.
Published: (2026) -
Pipeline Parallelism with Controllable Memory
by: Qi, Penghui, et al.
Published: (2024) -
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
by: Jin, Hongyi, et al.
Published: (2026) -
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
by: Huang, Mincong, et al.
Published: (2024)