HoliSDiP: Image Super-Resolution via Holistic Semantics and Diffusion Prior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tsao, Li-Yuan, Chen, Hao-Wei, Chung, Hao-Wei, Sun, Deqing, Lee, Chun-Yi, Chan, Kelvin C. K., Yang, Ming-Hsuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913589937307648
author Tsao, Li-Yuan
Chen, Hao-Wei
Chung, Hao-Wei
Sun, Deqing
Lee, Chun-Yi
Chan, Kelvin C. K.
Yang, Ming-Hsuan
author_facet Tsao, Li-Yuan
Chen, Hao-Wei
Chung, Hao-Wei
Sun, Deqing
Lee, Chun-Yi
Chan, Kelvin C. K.
Yang, Ming-Hsuan
contents Text-to-image diffusion models have emerged as powerful priors for real-world image super-resolution (Real-ISR). However, existing methods may produce unintended results due to noisy text prompts and their lack of spatial information. In this paper, we present HoliSDiP, a framework that leverages semantic segmentation to provide both precise textual and spatial guidance for diffusion-based Real-ISR. Our method employs semantic labels as concise text prompts while introducing dense semantic guidance through segmentation masks and our proposed Segmentation-CLIP Map. Extensive experiments demonstrate that HoliSDiP achieves significant improvement in image quality across various Real-ISR scenarios through reduced prompt noise and enhanced spatial control.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18662
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HoliSDiP: Image Super-Resolution via Holistic Semantics and Diffusion Prior
Tsao, Li-Yuan
Chen, Hao-Wei
Chung, Hao-Wei
Sun, Deqing
Lee, Chun-Yi
Chan, Kelvin C. K.
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
Text-to-image diffusion models have emerged as powerful priors for real-world image super-resolution (Real-ISR). However, existing methods may produce unintended results due to noisy text prompts and their lack of spatial information. In this paper, we present HoliSDiP, a framework that leverages semantic segmentation to provide both precise textual and spatial guidance for diffusion-based Real-ISR. Our method employs semantic labels as concise text prompts while introducing dense semantic guidance through segmentation masks and our proposed Segmentation-CLIP Map. Extensive experiments demonstrate that HoliSDiP achieves significant improvement in image quality across various Real-ISR scenarios through reduced prompt noise and enhanced spatial control.
title HoliSDiP: Image Super-Resolution via Holistic Semantics and Diffusion Prior
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.18662