Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Boqiang, Xie, Hongtao, Gao, Zuan, Wang, Yuxin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929337173803008
author Zhang, Boqiang
Xie, Hongtao
Gao, Zuan
Wang, Yuxin
author_facet Zhang, Boqiang
Xie, Hongtao
Gao, Zuan
Wang, Yuxin
contents Scene text images contain not only style information (font, background) but also content information (character, texture). Different scene text tasks need different information, but previous representation learning methods use tightly coupled features for all tasks, resulting in sub-optimal performance. We propose a Disentangled Representation Learning framework (DARLING) aimed at disentangling these two types of features for improved adaptability in better addressing various downstream tasks (choose what you really need). Specifically, we synthesize a dataset of image pairs with identical style but different content. Based on the dataset, we decouple the two types of features by the supervision design. Clearly, we directly split the visual representation into style and content features, the content features are supervised by a text recognition loss, while an alignment loss aligns the style features in the image pairs. Then, style features are employed in reconstructing the counterpart image via an image decoder with a prompt that indicates the counterpart's content. Such an operation effectively decouples the features based on their distinctive properties. To the best of our knowledge, this is the first time in the field of scene text that disentangles the inherent properties of the text images. Our method achieves state-of-the-art performance in Scene Text Recognition, Removal, and Editing.
format Preprint
id arxiv_https___arxiv_org_abs_2405_04377
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing
Zhang, Boqiang
Xie, Hongtao
Gao, Zuan
Wang, Yuxin
Computer Vision and Pattern Recognition
Scene text images contain not only style information (font, background) but also content information (character, texture). Different scene text tasks need different information, but previous representation learning methods use tightly coupled features for all tasks, resulting in sub-optimal performance. We propose a Disentangled Representation Learning framework (DARLING) aimed at disentangling these two types of features for improved adaptability in better addressing various downstream tasks (choose what you really need). Specifically, we synthesize a dataset of image pairs with identical style but different content. Based on the dataset, we decouple the two types of features by the supervision design. Clearly, we directly split the visual representation into style and content features, the content features are supervised by a text recognition loss, while an alignment loss aligns the style features in the image pairs. Then, style features are employed in reconstructing the counterpart image via an image decoder with a prompt that indicates the counterpart's content. Such an operation effectively decouples the features based on their distinctive properties. To the best of our knowledge, this is the first time in the field of scene text that disentangles the inherent properties of the text images. Our method achieves state-of-the-art performance in Scene Text Recognition, Removal, and Editing.
title Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.04377