Text-Aware Image Restoration with Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Min, Jaewon, Kim, Jin Hyeon, Cho, Paul Hyunbin, Lee, Jaeeun, Park, Jihye, Park, Minkyu, Kim, Sangpil, Park, Hyunhee, Kim, Seungryong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912461551042560
author Min, Jaewon
Kim, Jin Hyeon
Cho, Paul Hyunbin
Lee, Jaeeun
Park, Jihye
Park, Minkyu
Kim, Sangpil
Park, Hyunhee
Kim, Seungryong
author_facet Min, Jaewon
Kim, Jin Hyeon
Cho, Paul Hyunbin
Lee, Jaeeun
Park, Jihye
Park, Minkyu
Kim, Sangpil
Park, Hyunhee
Kim, Seungryong
contents Image restoration aims to recover degraded images. However, existing diffusion-based restoration methods, despite great success in natural image restoration, often struggle to faithfully reconstruct textual regions in degraded images. Those methods frequently generate plausible but incorrect text-like patterns, a phenomenon we refer to as text-image hallucination. In this paper, we introduce Text-Aware Image Restoration (TAIR), a novel restoration task that requires the simultaneous recovery of visual contents and textual fidelity. To tackle this task, we present SA-Text, a large-scale benchmark of 100K high-quality scene images densely annotated with diverse and complex text instances. Furthermore, we propose a multi-task diffusion framework, called TeReDiff, that integrates internal features from diffusion models into a text-spotting module, enabling both components to benefit from joint training. This allows for the extraction of rich text representations, which are utilized as prompts in subsequent denoising steps. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art restoration methods, achieving significant gains in text recognition accuracy. See our project page: https://cvlab-kaist.github.io/TAIR/
format Preprint
id arxiv_https___arxiv_org_abs_2506_09993
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text-Aware Image Restoration with Diffusion Models
Min, Jaewon
Kim, Jin Hyeon
Cho, Paul Hyunbin
Lee, Jaeeun
Park, Jihye
Park, Minkyu
Kim, Sangpil
Park, Hyunhee
Kim, Seungryong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Image restoration aims to recover degraded images. However, existing diffusion-based restoration methods, despite great success in natural image restoration, often struggle to faithfully reconstruct textual regions in degraded images. Those methods frequently generate plausible but incorrect text-like patterns, a phenomenon we refer to as text-image hallucination. In this paper, we introduce Text-Aware Image Restoration (TAIR), a novel restoration task that requires the simultaneous recovery of visual contents and textual fidelity. To tackle this task, we present SA-Text, a large-scale benchmark of 100K high-quality scene images densely annotated with diverse and complex text instances. Furthermore, we propose a multi-task diffusion framework, called TeReDiff, that integrates internal features from diffusion models into a text-spotting module, enabling both components to benefit from joint training. This allows for the extraction of rich text representations, which are utilized as prompts in subsequent denoising steps. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art restoration methods, achieving significant gains in text recognition accuracy. See our project page: https://cvlab-kaist.github.io/TAIR/
title Text-Aware Image Restoration with Diffusion Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.09993