Class-Aware Mask-Guided Feature Refinement for Scene Text Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Mingkun, Yang, Biao, Liao, Minghui, Zhu, Yingying, Bai, Xiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917594932445184
author Yang, Mingkun
Yang, Biao
Liao, Minghui
Zhu, Yingying
Bai, Xiang
author_facet Yang, Mingkun
Yang, Biao
Liao, Minghui
Zhu, Yingying
Bai, Xiang
contents Scene text recognition is a rapidly developing field that faces numerous challenges due to the complexity and diversity of scene text, including complex backgrounds, diverse fonts, flexible arrangements, and accidental occlusions. In this paper, we propose a novel approach called Class-Aware Mask-guided feature refinement (CAM) to address these challenges. Our approach introduces canonical class-aware glyph masks generated from a standard font to effectively suppress background and text style noise, thereby enhancing feature discrimination. Additionally, we design a feature alignment and fusion module to incorporate the canonical mask guidance for further feature refinement for text recognition. By enhancing the alignment between the canonical mask feature and the text feature, the module ensures more effective fusion, ultimately leading to improved recognition performance. We first evaluate CAM on six standard text recognition benchmarks to demonstrate its effectiveness. Furthermore, CAM exhibits superiority over the state-of-the-art method by an average performance gain of 4.1% across six more challenging datasets, despite utilizing a smaller model size. Our study highlights the importance of incorporating canonical mask guidance and aligned feature refinement techniques for robust scene text recognition. The code is available at https://github.com/MelosY/CAM.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13643
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Class-Aware Mask-Guided Feature Refinement for Scene Text Recognition
Yang, Mingkun
Yang, Biao
Liao, Minghui
Zhu, Yingying
Bai, Xiang
Computer Vision and Pattern Recognition
Scene text recognition is a rapidly developing field that faces numerous challenges due to the complexity and diversity of scene text, including complex backgrounds, diverse fonts, flexible arrangements, and accidental occlusions. In this paper, we propose a novel approach called Class-Aware Mask-guided feature refinement (CAM) to address these challenges. Our approach introduces canonical class-aware glyph masks generated from a standard font to effectively suppress background and text style noise, thereby enhancing feature discrimination. Additionally, we design a feature alignment and fusion module to incorporate the canonical mask guidance for further feature refinement for text recognition. By enhancing the alignment between the canonical mask feature and the text feature, the module ensures more effective fusion, ultimately leading to improved recognition performance. We first evaluate CAM on six standard text recognition benchmarks to demonstrate its effectiveness. Furthermore, CAM exhibits superiority over the state-of-the-art method by an average performance gain of 4.1% across six more challenging datasets, despite utilizing a smaller model size. Our study highlights the importance of incorporating canonical mask guidance and aligned feature refinement techniques for robust scene text recognition. The code is available at https://github.com/MelosY/CAM.
title Class-Aware Mask-Guided Feature Refinement for Scene Text Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.13643