Image Super-Resolution with Text Prompt Diffusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Zheng, Zhang, Yulun, Gu, Jinjin, Yuan, Xin, Kong, Linghe, Chen, Guihai, Yang, Xiaokang
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915190206889984
author Chen, Zheng
Zhang, Yulun
Gu, Jinjin
Yuan, Xin
Kong, Linghe
Chen, Guihai
Yang, Xiaokang
author_facet Chen, Zheng
Zhang, Yulun
Gu, Jinjin
Yuan, Xin
Kong, Linghe
Chen, Guihai
Yang, Xiaokang
contents Image super-resolution (SR) methods typically model degradation to improve reconstruction accuracy in complex and unknown degradation scenarios. However, extracting degradation information from low-resolution images is challenging, which limits the model performance. To boost image SR performance, one feasible approach is to introduce additional priors. Inspired by advancements in multi-modal methods and text prompt image processing, we introduce text prompts to image SR to provide degradation priors. Specifically, we first design a text-image generation pipeline to integrate text into the SR dataset through the text degradation representation and degradation model. By adopting a discrete design, the text representation is flexible and user-friendly. Meanwhile, we propose the PromptSR to realize the text prompt SR. The PromptSR leverages the latest multi-modal large language model (MLLM) to generate prompts from low-resolution images. It also utilizes the pre-trained language model (e.g., T5 or CLIP) to enhance text comprehension. We train the PromptSR on the text-image dataset. Extensive experiments indicate that introducing text prompts into SR, yields impressive results on both synthetic and real-world images. Code: https://github.com/zhengchen1999/PromptSR.
format Preprint
id arxiv_https___arxiv_org_abs_2311_14282
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Image Super-Resolution with Text Prompt Diffusion
Chen, Zheng
Zhang, Yulun
Gu, Jinjin
Yuan, Xin
Kong, Linghe
Chen, Guihai
Yang, Xiaokang
Computer Vision and Pattern Recognition
Image super-resolution (SR) methods typically model degradation to improve reconstruction accuracy in complex and unknown degradation scenarios. However, extracting degradation information from low-resolution images is challenging, which limits the model performance. To boost image SR performance, one feasible approach is to introduce additional priors. Inspired by advancements in multi-modal methods and text prompt image processing, we introduce text prompts to image SR to provide degradation priors. Specifically, we first design a text-image generation pipeline to integrate text into the SR dataset through the text degradation representation and degradation model. By adopting a discrete design, the text representation is flexible and user-friendly. Meanwhile, we propose the PromptSR to realize the text prompt SR. The PromptSR leverages the latest multi-modal large language model (MLLM) to generate prompts from low-resolution images. It also utilizes the pre-trained language model (e.g., T5 or CLIP) to enhance text comprehension. We train the PromptSR on the text-image dataset. Extensive experiments indicate that introducing text prompts into SR, yields impressive results on both synthetic and real-world images. Code: https://github.com/zhengchen1999/PromptSR.
title Image Super-Resolution with Text Prompt Diffusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.14282