A Wavelet Diffusion GAN for Image Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aloisi, Lorenzo, Sigillo, Luigi, Uncini, Aurelio, Comminiello, Danilo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913095826276352
author Aloisi, Lorenzo
Sigillo, Luigi
Uncini, Aurelio
Comminiello, Danilo
author_facet Aloisi, Lorenzo
Sigillo, Luigi
Uncini, Aurelio
Comminiello, Danilo
contents In recent years, diffusion models have emerged as a superior alternative to generative adversarial networks (GANs) for high-fidelity image generation, with wide applications in text-to-image generation, image-to-image translation, and super-resolution. However, their real-time feasibility is hindered by slow training and inference speeds. This study addresses this challenge by proposing a wavelet-based conditional Diffusion GAN scheme for Single-Image Super-Resolution (SISR). Our approach utilizes the diffusion GAN paradigm to reduce the timesteps required by the reverse diffusion process and the Discrete Wavelet Transform (DWT) to achieve dimensionality reduction, decreasing training and inference times significantly. The results of an experimental validation on the CelebA-HQ dataset confirm the effectiveness of our proposed scheme. Our approach outperforms other state-of-the-art methodologies successfully ensuring high-fidelity output while overcoming inherent drawbacks associated with diffusion models in time-sensitive applications. The code is available at https://www.github.com/aloilor/WaDiGAN-SR
format Preprint
id arxiv_https___arxiv_org_abs_2410_17966
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Wavelet Diffusion GAN for Image Super-Resolution
Aloisi, Lorenzo
Sigillo, Luigi
Uncini, Aurelio
Comminiello, Danilo
Image and Video Processing
Computer Vision and Pattern Recognition
In recent years, diffusion models have emerged as a superior alternative to generative adversarial networks (GANs) for high-fidelity image generation, with wide applications in text-to-image generation, image-to-image translation, and super-resolution. However, their real-time feasibility is hindered by slow training and inference speeds. This study addresses this challenge by proposing a wavelet-based conditional Diffusion GAN scheme for Single-Image Super-Resolution (SISR). Our approach utilizes the diffusion GAN paradigm to reduce the timesteps required by the reverse diffusion process and the Discrete Wavelet Transform (DWT) to achieve dimensionality reduction, decreasing training and inference times significantly. The results of an experimental validation on the CelebA-HQ dataset confirm the effectiveness of our proposed scheme. Our approach outperforms other state-of-the-art methodologies successfully ensuring high-fidelity output while overcoming inherent drawbacks associated with diffusion models in time-sensitive applications. The code is available at https://www.github.com/aloilor/WaDiGAN-SR
title A Wavelet Diffusion GAN for Image Super-Resolution
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.17966