KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kai, Wang, Peng, Bi, Sai, Zhang, Jianming, Xiong, Yuanjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911098172604416
author Zhang, Kai
Wang, Peng
Bi, Sai
Zhang, Jianming
Xiong, Yuanjun
author_facet Zhang, Kai
Wang, Peng
Bi, Sai
Zhang, Jianming
Xiong, Yuanjun
contents We present KnapFormer, an efficient and versatile framework to combine workload balancing and sequence parallelism in distributed training of Diffusion Transformers (DiT). KnapFormer builds on the insight that strong synergy exists between sequence parallelism and the need to address the significant token imbalance across ranks. This imbalance arises from variable-length text inputs and varying visual token counts in mixed-resolution and image-video joint training. KnapFormer redistributes tokens by first gathering sequence length metadata across all ranks in a balancing group and solving a global knapsack problem. The solver aims to minimize the variances of total workload per-GPU, while accounting for the effect of sequence parallelism. By integrating DeepSpeed-Ulysees-based sequence parallelism in the load-balancing decision process and utilizing a simple semi-empirical workload model, KnapFormers achieves minimal communication overhead and less than 1% workload discrepancy in real-world training workloads with sequence length varying from a few hundred to tens of thousands. It eliminates straggler effects and achieves 2x to 3x speedup when training state-of-the-art diffusion models like FLUX on mixed-resolution and image-video joint data corpora. We open-source the KnapFormer implementation at https://github.com/Kai-46/KnapFormer/
format Preprint
id arxiv_https___arxiv_org_abs_2508_06001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
Zhang, Kai
Wang, Peng
Bi, Sai
Zhang, Jianming
Xiong, Yuanjun
Distributed, Parallel, and Cluster Computing
Computer Vision and Pattern Recognition
We present KnapFormer, an efficient and versatile framework to combine workload balancing and sequence parallelism in distributed training of Diffusion Transformers (DiT). KnapFormer builds on the insight that strong synergy exists between sequence parallelism and the need to address the significant token imbalance across ranks. This imbalance arises from variable-length text inputs and varying visual token counts in mixed-resolution and image-video joint training. KnapFormer redistributes tokens by first gathering sequence length metadata across all ranks in a balancing group and solving a global knapsack problem. The solver aims to minimize the variances of total workload per-GPU, while accounting for the effect of sequence parallelism. By integrating DeepSpeed-Ulysees-based sequence parallelism in the load-balancing decision process and utilizing a simple semi-empirical workload model, KnapFormers achieves minimal communication overhead and less than 1% workload discrepancy in real-world training workloads with sequence length varying from a few hundred to tens of thousands. It eliminates straggler effects and achieves 2x to 3x speedup when training state-of-the-art diffusion models like FLUX on mixed-resolution and image-video joint data corpora. We open-source the KnapFormer implementation at https://github.com/Kai-46/KnapFormer/
title KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
topic Distributed, Parallel, and Cluster Computing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.06001