DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Zhenliang, Hu, Hanpeng, Chen, Xing, Jiang, Yimin, Song, Yixin, Mi, Zeyu, Zhu, Yibo, Jiang, Daxin, Xia, Yubin, Chen, Haibo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911536356786176
author Xue, Zhenliang
Hu, Hanpeng
Chen, Xing
Jiang, Yimin
Song, Yixin
Mi, Zeyu
Zhu, Yibo
Jiang, Daxin
Xia, Yubin
Chen, Haibo
author_facet Xue, Zhenliang
Hu, Hanpeng
Chen, Xing
Jiang, Yimin
Song, Yixin
Mi, Zeyu
Zhu, Yibo
Jiang, Daxin
Xia, Yubin
Chen, Haibo
contents Large multimodal models (LMMs) have demonstrated excellent capabilities in both understanding and generation tasks with various modalities. While these models can accept flexible combinations of input data, their training efficiency suffers from two major issues: pipeline stage imbalance caused by heterogeneous model architectures, and training data dynamicity stemming from the diversity of multimodal data. In this paper, we present DIP, a dynamic and modality-aware pipeline scheduling framework designed for LMM training. DIP tackles the challenge of dynamic imbalance via two key techniques: (1) separating computations of different modalities into dedicated pipeline segments to balance workloads within a continuous set of stages; (2) dynamically splitting input data into finer-grained, modality-specific sub-microbatches to balance workloads across these segments. By asynchronously generating pipeline schedules on idle CPU resources during training, DIP dynamically tailors stage executions to each input batch without stalling the training process. We validate DIP on a diverse set of five LMMs, ranging from 12B to 94B parameters and including vision-language and diffusion models. Experimental results show that our system achieves up to 97.3% higher throughput compared to state-of-the-art systems, demonstrating strong adaptability to fluctuating multimodal training workloads.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14145
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
Xue, Zhenliang
Hu, Hanpeng
Chen, Xing
Jiang, Yimin
Song, Yixin
Mi, Zeyu
Zhu, Yibo
Jiang, Daxin
Xia, Yubin
Chen, Haibo
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Large multimodal models (LMMs) have demonstrated excellent capabilities in both understanding and generation tasks with various modalities. While these models can accept flexible combinations of input data, their training efficiency suffers from two major issues: pipeline stage imbalance caused by heterogeneous model architectures, and training data dynamicity stemming from the diversity of multimodal data. In this paper, we present DIP, a dynamic and modality-aware pipeline scheduling framework designed for LMM training. DIP tackles the challenge of dynamic imbalance via two key techniques: (1) separating computations of different modalities into dedicated pipeline segments to balance workloads within a continuous set of stages; (2) dynamically splitting input data into finer-grained, modality-specific sub-microbatches to balance workloads across these segments. By asynchronously generating pipeline schedules on idle CPU resources during training, DIP dynamically tailors stage executions to each input batch without stalling the training process. We validate DIP on a diverse set of five LMMs, ranging from 12B to 94B parameters and including vision-language and diffusion models. Experimental results show that our system achieves up to 97.3% higher throughput compared to state-of-the-art systems, demonstrating strong adaptability to fluctuating multimodal training workloads.
title DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2504.14145