Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jang, Insu, Chowdhury, Mosharaf
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916053175500800
author Jang, Insu
Chowdhury, Mosharaf
author_facet Jang, Insu
Chowdhury, Mosharaf
contents Multimodal LLM datasets are inherently heterogeneous, with significant data variability. Although each modality exhibits independent variability, sample-level entanglement makes it difficult to balance workloads across both modalities and batches. We present Entrain, a distributed MLLM training framework that addresses both heterogeneity and variability in multimodal training workloads. Entrain challenges the intuition that dynamic data variability requires dynamic model parallelism by shifting the profiling paradigm from micro-level samples to macroscopic batches. We prove that a single, static model-parallel configuration suffices for optimal load balancing under this paradigm. At the microscopic scale, Entrain introduces a hierarchical microbatch assignment algorithm that defers excess workload within each iteration to stabilize variability across microbatches. Evaluations show that Entrain reduces workload variability across microbatches by up to 10.6$\times$, improving end-to-end training throughput by up to 1.40$\times$ over existing baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27918
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain
Jang, Insu
Chowdhury, Mosharaf
Distributed, Parallel, and Cluster Computing
Multimodal LLM datasets are inherently heterogeneous, with significant data variability. Although each modality exhibits independent variability, sample-level entanglement makes it difficult to balance workloads across both modalities and batches. We present Entrain, a distributed MLLM training framework that addresses both heterogeneity and variability in multimodal training workloads. Entrain challenges the intuition that dynamic data variability requires dynamic model parallelism by shifting the profiling paradigm from micro-level samples to macroscopic batches. We prove that a single, static model-parallel configuration suffices for optimal load balancing under this paradigm. At the microscopic scale, Entrain introduces a hierarchical microbatch assignment algorithm that defers excess workload within each iteration to stabilize variability across microbatches. Evaluations show that Entrain reduces workload variability across microbatches by up to 10.6$\times$, improving end-to-end training throughput by up to 1.40$\times$ over existing baselines.
title Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2605.27918