Saved in:
Bibliographic Details
Main Authors: Park, Dongwoo, Ko, Suk Pil
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.00410
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912302340505600
author Park, Dongwoo
Ko, Suk Pil
author_facet Park, Dongwoo
Ko, Suk Pil
contents Scene text image super-resolution (STISR) enhances the resolution and quality of low-resolution images. Unlike previous studies that treated scene text images as natural images, recent methods using a text prior (TP), extracted from a pre-trained text recognizer, have shown strong performance. However, two major issues emerge: (1) Explicit categorical priors, like TP, can negatively impact STISR if incorrect. We reveal that these explicit priors are unstable and propose replacing them with Non-CAtegorical Prior (NCAP) using penultimate layer representations. (2) Pre-trained recognizers used to generate TP struggle with low-resolution images. To address this, most studies jointly train the recognizer with the STISR network to bridge the domain gap between low- and high-resolution images, but this can cause an overconfidence phenomenon in the prior modality. We highlight this issue and propose a method to mitigate it by mixing hard and soft labels. Experiments on the TextZoom dataset demonstrate an improvement by 3.5%, while our method significantly enhances generalization performance by 14.8\% across four text recognition datasets. Our method generalizes to all TP-guided STISR networks.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00410
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NCAP: Scene Text Image Super-Resolution with Non-CAtegorical Prior
Park, Dongwoo
Ko, Suk Pil
Computer Vision and Pattern Recognition
Scene text image super-resolution (STISR) enhances the resolution and quality of low-resolution images. Unlike previous studies that treated scene text images as natural images, recent methods using a text prior (TP), extracted from a pre-trained text recognizer, have shown strong performance. However, two major issues emerge: (1) Explicit categorical priors, like TP, can negatively impact STISR if incorrect. We reveal that these explicit priors are unstable and propose replacing them with Non-CAtegorical Prior (NCAP) using penultimate layer representations. (2) Pre-trained recognizers used to generate TP struggle with low-resolution images. To address this, most studies jointly train the recognizer with the STISR network to bridge the domain gap between low- and high-resolution images, but this can cause an overconfidence phenomenon in the prior modality. We highlight this issue and propose a method to mitigate it by mixing hard and soft labels. Experiments on the TextZoom dataset demonstrate an improvement by 3.5%, while our method significantly enhances generalization performance by 14.8\% across four text recognition datasets. Our method generalizes to all TP-guided STISR networks.
title NCAP: Scene Text Image Super-Resolution with Non-CAtegorical Prior
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.00410