PiT: Progressive Diffusion Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jiafu, Wang, Yabiao, Li, Jian, Peng, Jinlong, Cao, Yun, Wang, Chengjie, Zhang, Jiangning
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913990105366528
author Wu, Jiafu
Wang, Yabiao
Li, Jian
Peng, Jinlong
Cao, Yun
Wang, Chengjie
Zhang, Jiangning
author_facet Wu, Jiafu
Wang, Yabiao
Li, Jian
Peng, Jinlong
Cao, Yun
Wang, Chengjie
Zhang, Jiangning
contents Diffusion Transformers (DiTs) achieve remarkable performance within image generation via the transformer architecture. Conventionally, DiTs are constructed by stacking serial isotropic global modeling transformers, which face significant quadratic computational cost. However, through empirical analysis, we find that DiTs do not rely as heavily on global information as previously believed. In fact, most layers exhibit significant redundancy in global computation. Additionally, conventional attention mechanisms suffer from low-frequency inertia, limiting their efficiency. To address these issues, we propose Pseudo Shifted Window Attention (PSWA), which fundamentally mitigates global attention redundancy. PSWA achieves moderate global-local information through window attention. It further utilizes a high-frequency bridging branch to simulate shifted window operations, which both enrich the high-frequency information and strengthen inter-window connections. Furthermore, we propose the Progressive Coverage Channel Allocation (PCCA) strategy that captures high-order attention without additional computational cost. Based on these innovations, we propose a series of Pseudo Progressive Diffusion Transformer (PiT). Our extensive experiments show their superior performance; for example, our proposed PiT-L achieves 54% FID improvement over DiT-XL/2 while using less computation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13219
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PiT: Progressive Diffusion Transformer
Wu, Jiafu
Wang, Yabiao
Li, Jian
Peng, Jinlong
Cao, Yun
Wang, Chengjie
Zhang, Jiangning
Computer Vision and Pattern Recognition
Diffusion Transformers (DiTs) achieve remarkable performance within image generation via the transformer architecture. Conventionally, DiTs are constructed by stacking serial isotropic global modeling transformers, which face significant quadratic computational cost. However, through empirical analysis, we find that DiTs do not rely as heavily on global information as previously believed. In fact, most layers exhibit significant redundancy in global computation. Additionally, conventional attention mechanisms suffer from low-frequency inertia, limiting their efficiency. To address these issues, we propose Pseudo Shifted Window Attention (PSWA), which fundamentally mitigates global attention redundancy. PSWA achieves moderate global-local information through window attention. It further utilizes a high-frequency bridging branch to simulate shifted window operations, which both enrich the high-frequency information and strengthen inter-window connections. Furthermore, we propose the Progressive Coverage Channel Allocation (PCCA) strategy that captures high-order attention without additional computational cost. Based on these innovations, we propose a series of Pseudo Progressive Diffusion Transformer (PiT). Our extensive experiments show their superior performance; for example, our proposed PiT-L achieves 54% FID improvement over DiT-XL/2 while using less computation.
title PiT: Progressive Diffusion Transformer
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.13219