Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Luo, Minxing, Fan, Linlong, Qiushi, Wang, Wu, Ge, Luo, Yiyan, Yu, Yuhang, Chen, Jinwei, Wang, Yaxing, Fan, Qingnan, Yang, Jian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915634348032000
author Luo, Minxing
Fan, Linlong
Qiushi, Wang
Wu, Ge
Luo, Yiyan
Yu, Yuhang
Chen, Jinwei
Wang, Yaxing
Fan, Qingnan
Yang, Jian
author_facet Luo, Minxing
Fan, Linlong
Qiushi, Wang
Wu, Ge
Luo, Yiyan
Yu, Yuhang
Chen, Jinwei
Wang, Yaxing
Fan, Qingnan
Yang, Jian
contents Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image quality and textual readability. To address this, we introduce TIGER (Text-Image Guided supEr-Resolution), a novel two-stage framework that breaks this trade-off through a "text-first, image-later" paradigm. TIGER explicitly decouples glyph restoration from image enhancement: it first reconstructs precise text structures and uses them to guide full-image super-resolution. This ensures high fidelity and readability. To support comprehensive training and evaluation, we present the UZ-ST (UltraZoom-Scene Text) dataset, the first Chinese scene text dataset with extreme zoom. Extensive experiments show TIGER achieves state-of-the-art performance, enhancing readability and image quality.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21590
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance
Luo, Minxing
Fan, Linlong
Qiushi, Wang
Wu, Ge
Luo, Yiyan
Yu, Yuhang
Chen, Jinwei
Wang, Yaxing
Fan, Qingnan
Yang, Jian
Computer Vision and Pattern Recognition
Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image quality and textual readability. To address this, we introduce TIGER (Text-Image Guided supEr-Resolution), a novel two-stage framework that breaks this trade-off through a "text-first, image-later" paradigm. TIGER explicitly decouples glyph restoration from image enhancement: it first reconstructs precise text structures and uses them to guide full-image super-resolution. This ensures high fidelity and readability. To support comprehensive training and evaluation, we present the UZ-ST (UltraZoom-Scene Text) dataset, the first Chinese scene text dataset with extreme zoom. Extensive experiments show TIGER achieves state-of-the-art performance, enhancing readability and image quality.
title Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.21590