Effective Diffusion Transformer Architecture for Image Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Kun, Yu, Lei, Tu, Zhijun, He, Xiao, Chen, Liyu, Guo, Yong, Zhu, Mingrui, Wang, Nannan, Gao, Xinbo, Hu, Jie
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909328972185600
author Cheng, Kun
Yu, Lei
Tu, Zhijun
He, Xiao
Chen, Liyu
Guo, Yong
Zhu, Mingrui
Wang, Nannan
Gao, Xinbo
Hu, Jie
author_facet Cheng, Kun
Yu, Lei
Tu, Zhijun
He, Xiao
Chen, Liyu
Guo, Yong
Zhu, Mingrui
Wang, Nannan
Gao, Xinbo
Hu, Jie
contents Recent advances indicate that diffusion models hold great promise in image super-resolution. While the latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image generation. In this work, we design an effective diffusion transformer for image super-resolution (DiT-SR) that achieves the visual quality of prior-based methods, but through a training-from-scratch manner. In practice, DiT-SR leverages an overall U-shaped architecture, and adopts a uniform isotropic design for all the transformer blocks across different stages. The former facilitates multi-scale hierarchical feature extraction, while the latter reallocates the computational resources to critical layers to further enhance performance. Moreover, we thoroughly analyze the limitation of the widely used AdaLN, and present a frequency-adaptive time-step conditioning module, enhancing the model's capacity to process distinct frequency information at different time steps. Extensive experiments demonstrate that DiT-SR outperforms the existing training-from-scratch diffusion-based SR methods significantly, and even beats some of the prior-based methods on pretrained Stable Diffusion, proving the superiority of diffusion transformer in image super-resolution.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19589
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Effective Diffusion Transformer Architecture for Image Super-Resolution
Cheng, Kun
Yu, Lei
Tu, Zhijun
He, Xiao
Chen, Liyu
Guo, Yong
Zhu, Mingrui
Wang, Nannan
Gao, Xinbo
Hu, Jie
Computer Vision and Pattern Recognition
Recent advances indicate that diffusion models hold great promise in image super-resolution. While the latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image generation. In this work, we design an effective diffusion transformer for image super-resolution (DiT-SR) that achieves the visual quality of prior-based methods, but through a training-from-scratch manner. In practice, DiT-SR leverages an overall U-shaped architecture, and adopts a uniform isotropic design for all the transformer blocks across different stages. The former facilitates multi-scale hierarchical feature extraction, while the latter reallocates the computational resources to critical layers to further enhance performance. Moreover, we thoroughly analyze the limitation of the widely used AdaLN, and present a frequency-adaptive time-step conditioning module, enhancing the model's capacity to process distinct frequency information at different time steps. Extensive experiments demonstrate that DiT-SR outperforms the existing training-from-scratch diffusion-based SR methods significantly, and even beats some of the prior-based methods on pretrained Stable Diffusion, proving the superiority of diffusion transformer in image super-resolution.
title Effective Diffusion Transformer Architecture for Image Super-Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.19589