U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Yuchuan, Tu, Zhijun, Chen, Hanting, Hu, Jie, Xu, Chao, Wang, Yunhe
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917822740824064
author Tian, Yuchuan
Tu, Zhijun
Chen, Hanting
Hu, Jie
Xu, Chao
Wang, Yunhe
author_facet Tian, Yuchuan
Tu, Zhijun
Chen, Hanting
Hu, Jie
Xu, Chao
Wang, Yunhe
contents Diffusion Transformers (DiTs) introduce the transformer architecture to diffusion tasks for latent-space image generation. With an isotropic architecture that chains a series of transformer blocks, DiTs demonstrate competitive performance and good scalability; but meanwhile, the abandonment of U-Net by DiTs and their following improvements is worth rethinking. To this end, we conduct a simple toy experiment by comparing a U-Net architectured DiT with an isotropic one. It turns out that the U-Net architecture only gain a slight advantage amid the U-Net inductive bias, indicating potential redundancies within the U-Net-style DiT. Inspired by the discovery that U-Net backbone features are low-frequency-dominated, we perform token downsampling on the query-key-value tuple for self-attention that bring further improvements despite a considerable amount of reduction in computation. Based on self-attention with downsampled tokens, we propose a series of U-shaped DiTs (U-DiTs) in the paper and conduct extensive experiments to demonstrate the extraordinary performance of U-DiT models. The proposed U-DiT could outperform DiT-XL/2 with only 1/6 of its computation cost. Codes are available at https://github.com/YuchuanTian/U-DiT.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02730
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
Tian, Yuchuan
Tu, Zhijun
Chen, Hanting
Hu, Jie
Xu, Chao
Wang, Yunhe
Computer Vision and Pattern Recognition
Diffusion Transformers (DiTs) introduce the transformer architecture to diffusion tasks for latent-space image generation. With an isotropic architecture that chains a series of transformer blocks, DiTs demonstrate competitive performance and good scalability; but meanwhile, the abandonment of U-Net by DiTs and their following improvements is worth rethinking. To this end, we conduct a simple toy experiment by comparing a U-Net architectured DiT with an isotropic one. It turns out that the U-Net architecture only gain a slight advantage amid the U-Net inductive bias, indicating potential redundancies within the U-Net-style DiT. Inspired by the discovery that U-Net backbone features are low-frequency-dominated, we perform token downsampling on the query-key-value tuple for self-attention that bring further improvements despite a considerable amount of reduction in computation. Based on self-attention with downsampled tokens, we propose a series of U-shaped DiTs (U-DiTs) in the paper and conduct extensive experiments to demonstrate the extraordinary performance of U-DiT models. The proposed U-DiT could outperform DiT-XL/2 with only 1/6 of its computation cost. Codes are available at https://github.com/YuchuanTian/U-DiT.
title U-DiTs: Downsample Tokens in U-Shaped Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.02730