Boosting Diffusion-Based Text Image Super-Resolution Model Towards Generalized Real-World Scenarios

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Chenglu, Xu, Xiaogang, Ding, Ganggui, Zhang, Yunke, Li, Wenbo, Xu, Jiarong, Wu, Qingbiao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918058554032128
author Pan, Chenglu
Xu, Xiaogang
Ding, Ganggui
Zhang, Yunke
Li, Wenbo
Xu, Jiarong
Wu, Qingbiao
author_facet Pan, Chenglu
Xu, Xiaogang
Ding, Ganggui
Zhang, Yunke
Li, Wenbo
Xu, Jiarong
Wu, Qingbiao
contents Restoring low-resolution text images presents a significant challenge, as it requires maintaining both the fidelity and stylistic realism of the text in restored images. Existing text image restoration methods often fall short in hard situations, as the traditional super-resolution models cannot guarantee clarity, while diffusion-based methods fail to maintain fidelity. In this paper, we introduce a novel framework aimed at improving the generalization ability of diffusion models for text image super-resolution (SR), especially promoting fidelity. First, we propose a progressive data sampling strategy that incorporates diverse image types at different stages of training, stabilizing the convergence and improving the generalization. For the network architecture, we leverage a pre-trained SR prior to provide robust spatial reasoning capabilities, enhancing the model's ability to preserve textual information. Additionally, we employ a cross-attention mechanism to better integrate textual priors. To further reduce errors in textual priors, we utilize confidence scores to dynamically adjust the importance of textual features during training. Extensive experiments on real-world datasets demonstrate that our approach not only produces text images with more realistic visual appearances but also improves the accuracy of text structure.
format Preprint
id arxiv_https___arxiv_org_abs_2503_07232
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Boosting Diffusion-Based Text Image Super-Resolution Model Towards Generalized Real-World Scenarios
Pan, Chenglu
Xu, Xiaogang
Ding, Ganggui
Zhang, Yunke
Li, Wenbo
Xu, Jiarong
Wu, Qingbiao
Computer Vision and Pattern Recognition
Restoring low-resolution text images presents a significant challenge, as it requires maintaining both the fidelity and stylistic realism of the text in restored images. Existing text image restoration methods often fall short in hard situations, as the traditional super-resolution models cannot guarantee clarity, while diffusion-based methods fail to maintain fidelity. In this paper, we introduce a novel framework aimed at improving the generalization ability of diffusion models for text image super-resolution (SR), especially promoting fidelity. First, we propose a progressive data sampling strategy that incorporates diverse image types at different stages of training, stabilizing the convergence and improving the generalization. For the network architecture, we leverage a pre-trained SR prior to provide robust spatial reasoning capabilities, enhancing the model's ability to preserve textual information. Additionally, we employ a cross-attention mechanism to better integrate textual priors. To further reduce errors in textual priors, we utilize confidence scores to dynamically adjust the importance of textual features during training. Extensive experiments on real-world datasets demonstrate that our approach not only produces text images with more realistic visual appearances but also improves the accuracy of text structure.
title Boosting Diffusion-Based Text Image Super-Resolution Model Towards Generalized Real-World Scenarios
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.07232