DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ning, Mang, Li, Mingxiao, Su, Jianlin, Jia, Haozhe, Liu, Lanmiao, Beneš, Martin, Chen, Wenshuo, Salah, Albert Ali, Ertugrul, Itir Onal
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915313414569984
author Ning, Mang
Li, Mingxiao
Su, Jianlin
Jia, Haozhe
Liu, Lanmiao
Beneš, Martin
Chen, Wenshuo
Salah, Albert Ali
Ertugrul, Itir Onal
author_facet Ning, Mang
Li, Mingxiao
Su, Jianlin
Jia, Haozhe
Liu, Lanmiao
Beneš, Martin
Chen, Wenshuo
Salah, Albert Ali
Ertugrul, Itir Onal
contents This paper explores image modeling from the frequency space and introduces DCTdiff, an end-to-end diffusion generative paradigm that efficiently models images in the discrete cosine transform (DCT) space. We investigate the design space of DCTdiff and reveal the key design factors. Experiments on different frameworks (UViT, DiT), generation tasks, and various diffusion samplers demonstrate that DCTdiff outperforms pixel-based diffusion models regarding generative quality and training efficiency. Remarkably, DCTdiff can seamlessly scale up to 512$\times$512 resolution without using the latent diffusion paradigm and beats latent diffusion (using SD-VAE) with only 1/4 training cost. Finally, we illustrate several intriguing properties of DCT image modeling. For example, we provide a theoretical proof of why 'image diffusion can be seen as spectral autoregression', bridging the gap between diffusion and autoregressive models. The effectiveness of DCTdiff and the introduced properties suggest a promising direction for image modeling in the frequency space. The code is https://github.com/forever208/DCTdiff.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15032
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space
Ning, Mang
Li, Mingxiao
Su, Jianlin
Jia, Haozhe
Liu, Lanmiao
Beneš, Martin
Chen, Wenshuo
Salah, Albert Ali
Ertugrul, Itir Onal
Computer Vision and Pattern Recognition
Machine Learning
Image and Video Processing
This paper explores image modeling from the frequency space and introduces DCTdiff, an end-to-end diffusion generative paradigm that efficiently models images in the discrete cosine transform (DCT) space. We investigate the design space of DCTdiff and reveal the key design factors. Experiments on different frameworks (UViT, DiT), generation tasks, and various diffusion samplers demonstrate that DCTdiff outperforms pixel-based diffusion models regarding generative quality and training efficiency. Remarkably, DCTdiff can seamlessly scale up to 512$\times$512 resolution without using the latent diffusion paradigm and beats latent diffusion (using SD-VAE) with only 1/4 training cost. Finally, we illustrate several intriguing properties of DCT image modeling. For example, we provide a theoretical proof of why 'image diffusion can be seen as spectral autoregression', bridging the gap between diffusion and autoregressive models. The effectiveness of DCTdiff and the introduced properties suggest a promising direction for image modeling in the frequency space. The code is https://github.com/forever208/DCTdiff.
title DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT Space
topic Computer Vision and Pattern Recognition
Machine Learning
Image and Video Processing
url https://arxiv.org/abs/2412.15032