Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yifei, Liu, Chang, Wei, Jin, Yang, Xiaomeng, Zhou, Yu, Ma, Can, Ji, Xiangyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910890307092480
author Zhang, Yifei
Liu, Chang
Wei, Jin
Yang, Xiaomeng
Zhou, Yu
Ma, Can
Ji, Xiangyang
author_facet Zhang, Yifei
Liu, Chang
Wei, Jin
Yang, Xiaomeng
Zhou, Yu
Ma, Can
Ji, Xiangyang
contents Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic elements. In scenarios with degraded visual quality, linguistic patterns serve as crucial supplements for comprehension, highlighting the necessity of integrating both aspects for robust scene text recognition (STR). Contemporary STR approaches often use language models or semantic reasoning modules to capture linguistic features, typically requiring large-scale annotated datasets. Self-supervised learning, which lacks annotations, presents challenges in disentangling linguistic features related to the global context. Typically, sequence contrastive learning emphasizes the alignment of local features, while masked image modeling (MIM) tends to exploit local structures to reconstruct visual patterns, resulting in limited linguistic knowledge. In this paper, we propose a Linguistics-aware Masked Image Modeling (LMIM) approach, which channels the linguistic information into the decoding process of MIM through a separate branch. Specifically, we design a linguistics alignment module to extract vision-independent features as linguistic guidance using inputs with different visual appearances. As features extend beyond mere visual structures, LMIM must consider the global context to achieve reconstruction. Extensive experiments on various benchmarks quantitatively demonstrate our state-of-the-art performance, and attention visualizations qualitatively show the simultaneous capture of both visual and linguistic information.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18746
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
Zhang, Yifei
Liu, Chang
Wei, Jin
Yang, Xiaomeng
Zhou, Yu
Ma, Can
Ji, Xiangyang
Computer Vision and Pattern Recognition
Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates contextual and semantic elements. In scenarios with degraded visual quality, linguistic patterns serve as crucial supplements for comprehension, highlighting the necessity of integrating both aspects for robust scene text recognition (STR). Contemporary STR approaches often use language models or semantic reasoning modules to capture linguistic features, typically requiring large-scale annotated datasets. Self-supervised learning, which lacks annotations, presents challenges in disentangling linguistic features related to the global context. Typically, sequence contrastive learning emphasizes the alignment of local features, while masked image modeling (MIM) tends to exploit local structures to reconstruct visual patterns, resulting in limited linguistic knowledge. In this paper, we propose a Linguistics-aware Masked Image Modeling (LMIM) approach, which channels the linguistic information into the decoding process of MIM through a separate branch. Specifically, we design a linguistics alignment module to extract vision-independent features as linguistic guidance using inputs with different visual appearances. As features extend beyond mere visual structures, LMIM must consider the global context to achieve reconstruction. Extensive experiments on various benchmarks quantitatively demonstrate our state-of-the-art performance, and attention visualizations qualitatively show the simultaneous capture of both visual and linguistic information.
title Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.18746