AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Guo, Yucheng, Guo, Yongjian, Guan, Zhong, Sun, Haoran, Huang, Wen, Xu, Wanting, Long, Jing, Di, Shuai, Xiong, Junwu
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918509027524608
author Guo, Yucheng
Guo, Yongjian
Guan, Zhong
Sun, Haoran
Huang, Wen
Xu, Wanting
Long, Jing
Di, Shuai
Xiong, Junwu
author_facet Guo, Yucheng
Guo, Yongjian
Guan, Zhong
Sun, Haoran
Huang, Wen
Xu, Wanting
Long, Jing
Di, Shuai
Xiong, Junwu
contents In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant computational challenges due to the extreme variance in sequence lengths within mixed-mode datasets. Existing bucket-based data loading strategies typically rely on "equal token length" constraints. This approach fails to account for the quadratic complexity of self-attention mechanisms, leading to severe load imbalance and underutilization of GPU resources. This paper proposes \textit{AdaptiveLoad}, an integrated optimization framework consisting of two core components: (1) A dual-constraint adaptive load balancing system, which eliminates long-sequence bottlenecks by simultaneously limiting memory consumption and computational load ($B \times S^p \le M_{\text{comp}}$); (2) A fused LayerNorm-Modulate CUDA kernel, which utilizes a D-tile coalesced reduction strategy to increase throughput and alleviate memory pressure. Experimental results on the Wan 2.1 world model demonstrate that our method reduces the computational imbalance rate from 39\% to 18.9\%, improves peak VRAM utilization efficiency by 22.7\%, and achieves an overall training throughput increase of 27.2\%.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17923
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training
Guo, Yucheng
Guo, Yongjian
Guan, Zhong
Sun, Haoran
Huang, Wen
Xu, Wanting
Long, Jing
Di, Shuai
Xiong, Junwu
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant computational challenges due to the extreme variance in sequence lengths within mixed-mode datasets. Existing bucket-based data loading strategies typically rely on "equal token length" constraints. This approach fails to account for the quadratic complexity of self-attention mechanisms, leading to severe load imbalance and underutilization of GPU resources. This paper proposes \textit{AdaptiveLoad}, an integrated optimization framework consisting of two core components: (1) A dual-constraint adaptive load balancing system, which eliminates long-sequence bottlenecks by simultaneously limiting memory consumption and computational load ($B \times S^p \le M_{\text{comp}}$); (2) A fused LayerNorm-Modulate CUDA kernel, which utilizes a D-tile coalesced reduction strategy to increase throughput and alleviate memory pressure. Experimental results on the Wan 2.1 world model demonstrate that our method reduces the computational imbalance rate from 39\% to 18.9\%, improves peak VRAM utilization efficiency by 22.7\%, and achieves an overall training throughput increase of 27.2\%.
title AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.17923