Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Haoze, Li, Wenbo, Liu, Jiayue, Zhou, Kaiwen, Chen, Yongqiang, Guo, Yong, Li, Yanwei, Pei, Renjing, Peng, Long, Yang, Yujiu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910730624696320
author Sun, Haoze
Li, Wenbo
Liu, Jiayue
Zhou, Kaiwen
Chen, Yongqiang
Guo, Yong
Li, Yanwei
Pei, Renjing
Peng, Long
Yang, Yujiu
author_facet Sun, Haoze
Li, Wenbo
Liu, Jiayue
Zhou, Kaiwen
Chen, Yongqiang
Guo, Yong
Li, Yanwei
Pei, Renjing
Peng, Long
Yang, Yujiu
contents Generalization has long been a central challenge in real-world image restoration. While recent diffusion-based restoration methods, which leverage generative priors from text-to-image models, have made progress in recovering more realistic details, they still encounter "generative capability deactivation" when applied to out-of-distribution real-world data. To address this, we propose using text as an auxiliary invariant representation to reactivate the generative capabilities of these models. We begin by identifying two key properties of text input: richness and relevance, and examine their respective influence on model performance. Building on these insights, we introduce Res-Captioner, a module that generates enhanced textual descriptions tailored to image content and degradation levels, effectively mitigating response failures. Additionally, we present RealIR, a new benchmark designed to capture diverse real-world scenarios. Extensive experiments demonstrate that Res-Captioner significantly enhances the generalization abilities of diffusion-based restoration models, while remaining fully plug-and-play.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00878
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration
Sun, Haoze
Li, Wenbo
Liu, Jiayue
Zhou, Kaiwen
Chen, Yongqiang
Guo, Yong
Li, Yanwei
Pei, Renjing
Peng, Long
Yang, Yujiu
Computer Vision and Pattern Recognition
Generalization has long been a central challenge in real-world image restoration. While recent diffusion-based restoration methods, which leverage generative priors from text-to-image models, have made progress in recovering more realistic details, they still encounter "generative capability deactivation" when applied to out-of-distribution real-world data. To address this, we propose using text as an auxiliary invariant representation to reactivate the generative capabilities of these models. We begin by identifying two key properties of text input: richness and relevance, and examine their respective influence on model performance. Building on these insights, we introduce Res-Captioner, a module that generates enhanced textual descriptions tailored to image content and degradation levels, effectively mitigating response failures. Additionally, we present RealIR, a new benchmark designed to capture diverse real-world scenarios. Extensive experiments demonstrate that Res-Captioner significantly enhances the generalization abilities of diffusion-based restoration models, while remaining fully plug-and-play.
title Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00878