ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wan, Yuhao, Jiang, Peng-Tao, Hou, Qibin, Zhang, Hao, Chen, Jinwei, Cheng, Ming-Ming, Li, Bo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909560106647552
author Wan, Yuhao
Jiang, Peng-Tao
Hou, Qibin
Zhang, Hao
Chen, Jinwei
Cheng, Ming-Ming
Li, Bo
author_facet Wan, Yuhao
Jiang, Peng-Tao
Hou, Qibin
Zhang, Hao
Chen, Jinwei
Cheng, Ming-Ming
Li, Bo
contents We present ControlSR, a new method that can tame Diffusion Models for consistent real-world image super-resolution (Real-ISR). Previous Real-ISR models mostly focus on how to activate more generative priors of text-to-image diffusion models to make the output high-resolution (HR) images look better. However, since these methods rely too much on the generative priors, the content of the output images is often inconsistent with the input LR ones. To mitigate the above issue, in this work, we tame Diffusion Models by effectively utilizing LR information to impose stronger constraints on the control signals from ControlNet in the latent space. We show that our method can produce higher-quality control signals, which enables the super-resolution results to be more consistent with the LR image and leads to clearer visual results. In addition, we also propose an inference strategy that imposes constraints in the latent space using LR information, allowing for the simultaneous improvement of fidelity and generative ability. Experiments demonstrate that our model can achieve better performance across multiple metrics on several test sets and generate more consistent SR results with LR images than existing methods. Our code is available at https://github.com/HVision-NKU/ControlSR.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14279
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution
Wan, Yuhao
Jiang, Peng-Tao
Hou, Qibin
Zhang, Hao
Chen, Jinwei
Cheng, Ming-Ming
Li, Bo
Computer Vision and Pattern Recognition
We present ControlSR, a new method that can tame Diffusion Models for consistent real-world image super-resolution (Real-ISR). Previous Real-ISR models mostly focus on how to activate more generative priors of text-to-image diffusion models to make the output high-resolution (HR) images look better. However, since these methods rely too much on the generative priors, the content of the output images is often inconsistent with the input LR ones. To mitigate the above issue, in this work, we tame Diffusion Models by effectively utilizing LR information to impose stronger constraints on the control signals from ControlNet in the latent space. We show that our method can produce higher-quality control signals, which enables the super-resolution results to be more consistent with the LR image and leads to clearer visual results. In addition, we also propose an inference strategy that imposes constraints in the latent space using LR information, allowing for the simultaneous improvement of fidelity and generative ability. Experiments demonstrate that our model can achieve better performance across multiple metrics on several test sets and generate more consistent SR results with LR images than existing methods. Our code is available at https://github.com/HVision-NKU/ControlSR.
title ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.14279