Aggregated Text Transformer for Scene Text Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Zhao, Du, Xiangcheng, Zheng, Yingbin, Jin, Cheng
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910470943801344
author Zhou, Zhao
Du, Xiangcheng
Zheng, Yingbin
Jin, Cheng
author_facet Zhou, Zhao
Du, Xiangcheng
Zheng, Yingbin
Jin, Cheng
contents This paper explores the multi-scale aggregation strategy for scene text detection in natural images. We present the Aggregated Text TRansformer(ATTR), which is designed to represent texts in scene images with a multi-scale self-attention mechanism. Starting from the image pyramid with multiple resolutions, the features are first extracted at different scales with shared weight and then fed into an encoder-decoder architecture of Transformer. The multi-scale image representations are robust and contain rich information on text contents of various sizes. The text Transformer aggregates these features to learn the interaction across different scales and improve text representation. The proposed method detects scene texts by representing each text instance as an individual binary mask, which is tolerant of curve texts and regions with dense instances. Extensive experiments on public scene text detection datasets demonstrate the effectiveness of the proposed framework.
format Preprint
id arxiv_https___arxiv_org_abs_2211_13984
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Aggregated Text Transformer for Scene Text Detection
Zhou, Zhao
Du, Xiangcheng
Zheng, Yingbin
Jin, Cheng
Computer Vision and Pattern Recognition
This paper explores the multi-scale aggregation strategy for scene text detection in natural images. We present the Aggregated Text TRansformer(ATTR), which is designed to represent texts in scene images with a multi-scale self-attention mechanism. Starting from the image pyramid with multiple resolutions, the features are first extracted at different scales with shared weight and then fed into an encoder-decoder architecture of Transformer. The multi-scale image representations are robust and contain rich information on text contents of various sizes. The text Transformer aggregates these features to learn the interaction across different scales and improve text representation. The proposed method detects scene texts by representing each text instance as an individual binary mask, which is tolerant of curve texts and regions with dense instances. Extensive experiments on public scene text detection datasets demonstrate the effectiveness of the proposed framework.
title Aggregated Text Transformer for Scene Text Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2211.13984