EAFormer: Scene Text Segmentation with Edge-Aware Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Haiyang, Fu, Teng, Li, Bin, Xue, Xiangyang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917731763224576
author Yu, Haiyang
Fu, Teng
Li, Bin
Xue, Xiangyang
author_facet Yu, Haiyang
Fu, Teng
Li, Bin
Xue, Xiangyang
contents Scene text segmentation aims at cropping texts from scene images, which is usually used to help generative models edit or remove texts. The existing text segmentation methods tend to involve various text-related supervisions for better performance. However, most of them ignore the importance of text edges, which are significant for downstream applications. In this paper, we propose Edge-Aware Transformers, termed EAFormer, to segment texts more accurately, especially at the edge of texts. Specifically, we first design a text edge extractor to detect edges and filter out edges of non-text areas. Then, we propose an edge-guided encoder to make the model focus more on text edges. Finally, an MLP-based decoder is employed to predict text masks. We have conducted extensive experiments on commonly-used benchmarks to verify the effectiveness of EAFormer. The experimental results demonstrate that the proposed method can perform better than previous methods, especially on the segmentation of text edges. Considering that the annotations of several benchmarks (e.g., COCO_TS and MLT_S) are not accurate enough to fairly evaluate our methods, we have relabeled these datasets. Through experiments, we observe that our method can achieve a higher performance improvement when more accurate annotations are used for training.
format Preprint
id arxiv_https___arxiv_org_abs_2407_17020
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EAFormer: Scene Text Segmentation with Edge-Aware Transformers
Yu, Haiyang
Fu, Teng
Li, Bin
Xue, Xiangyang
Computer Vision and Pattern Recognition
Scene text segmentation aims at cropping texts from scene images, which is usually used to help generative models edit or remove texts. The existing text segmentation methods tend to involve various text-related supervisions for better performance. However, most of them ignore the importance of text edges, which are significant for downstream applications. In this paper, we propose Edge-Aware Transformers, termed EAFormer, to segment texts more accurately, especially at the edge of texts. Specifically, we first design a text edge extractor to detect edges and filter out edges of non-text areas. Then, we propose an edge-guided encoder to make the model focus more on text edges. Finally, an MLP-based decoder is employed to predict text masks. We have conducted extensive experiments on commonly-used benchmarks to verify the effectiveness of EAFormer. The experimental results demonstrate that the proposed method can perform better than previous methods, especially on the segmentation of text edges. Considering that the annotations of several benchmarks (e.g., COCO_TS and MLT_S) are not accurate enough to fairly evaluate our methods, we have relabeled these datasets. Through experiments, we observe that our method can achieve a higher performance improvement when more accurate annotations are used for training.
title EAFormer: Scene Text Segmentation with Edge-Aware Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.17020