DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Wangbo, Han, Yizeng, Tang, Jiasheng, Wang, Kai, Luo, Hao, Song, Yibing, Huang, Gao, Wang, Fan, You, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909989626445824
author Zhao, Wangbo
Han, Yizeng
Tang, Jiasheng
Wang, Kai
Luo, Hao
Song, Yibing
Huang, Gao
Wang, Fan
You, Yang
author_facet Zhao, Wangbo
Han, Yizeng
Tang, Jiasheng
Wang, Kai
Luo, Hao
Song, Yibing
Huang, Gao
Wang, Fan
You, Yang
contents Diffusion Transformer (DiT), an emerging diffusion model for visual generation, has demonstrated superior performance but suffers from substantial computational costs. Our investigations reveal that these costs primarily stem from the static inference paradigm, which inevitably introduces redundant computation in certain diffusion timesteps and spatial regions. To overcome this inefficiency, we propose Dynamic Diffusion Transformer (DyDiT), an architecture that dynamically adjusts its computation along both timestep and spatial dimensions. Building on these designs, we present an extended version, DyDiT++, with improvements in three key aspects. First, it extends the generation mechanism of DyDiT beyond diffusion to flow matching, demonstrating that our method can also accelerate flow-matching-based generation, enhancing its versatility. Furthermore, we enhance DyDiT to tackle more complex visual generation tasks, including video generation and text-to-image generation, thereby broadening its real-world applications. Finally, to address the high cost of full fine-tuning and democratize technology access, we investigate the feasibility of training DyDiT in a parameter-efficient manner and introduce timestep-based dynamic LoRA (TD-LoRA). Extensive experiments on diverse visual generation models, including DiT, SiT, Latte, and FLUX, demonstrate the effectiveness of DyDiT++. Remarkably, with <3% additional fine-tuning iterations, our approach reduces the FLOPs of DiT-XL by 51%, yielding 1.73x realistic speedup on hardware, and achieves a competitive FID score of 2.07 on ImageNet. The code is available at https://github.com/alibaba-damo-academy/DyDiT.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
Zhao, Wangbo
Han, Yizeng
Tang, Jiasheng
Wang, Kai
Luo, Hao
Song, Yibing
Huang, Gao
Wang, Fan
You, Yang
Computer Vision and Pattern Recognition
Diffusion Transformer (DiT), an emerging diffusion model for visual generation, has demonstrated superior performance but suffers from substantial computational costs. Our investigations reveal that these costs primarily stem from the static inference paradigm, which inevitably introduces redundant computation in certain diffusion timesteps and spatial regions. To overcome this inefficiency, we propose Dynamic Diffusion Transformer (DyDiT), an architecture that dynamically adjusts its computation along both timestep and spatial dimensions. Building on these designs, we present an extended version, DyDiT++, with improvements in three key aspects. First, it extends the generation mechanism of DyDiT beyond diffusion to flow matching, demonstrating that our method can also accelerate flow-matching-based generation, enhancing its versatility. Furthermore, we enhance DyDiT to tackle more complex visual generation tasks, including video generation and text-to-image generation, thereby broadening its real-world applications. Finally, to address the high cost of full fine-tuning and democratize technology access, we investigate the feasibility of training DyDiT in a parameter-efficient manner and introduce timestep-based dynamic LoRA (TD-LoRA). Extensive experiments on diverse visual generation models, including DiT, SiT, Latte, and FLUX, demonstrate the effectiveness of DyDiT++. Remarkably, with <3% additional fine-tuning iterations, our approach reduces the FLOPs of DiT-XL by 51%, yielding 1.73x realistic speedup on hardware, and achieves a competitive FID score of 2.07 on ImageNet. The code is available at https://github.com/alibaba-damo-academy/DyDiT.
title DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.06803