TextMastero: Mastering High-Quality Scene Text Editing in Diverse Languages and Styles

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Tong, Qu, Xiaochao, Liu, Ting
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910571053449216
author Wang, Tong
Qu, Xiaochao
Liu, Ting
author_facet Wang, Tong
Qu, Xiaochao
Liu, Ting
contents Scene text editing aims to modify texts on images while maintaining the style of newly generated text similar to the original. Given an image, a target area, and target text, the task produces an output image with the target text in the selected area, replacing the original. This task has been studied extensively, with initial success using Generative Adversarial Networks (GANs) to balance text fidelity and style similarity. However, GAN-based methods struggled with complex backgrounds or text styles. Recent works leverage diffusion models, showing improved results, yet still face challenges, especially with non-Latin languages like CJK characters (Chinese, Japanese, Korean) that have complex glyphs, often producing inaccurate or unrecognizable characters. To address these issues, we present \emph{TextMastero} - a carefully designed multilingual scene text editing architecture based on latent diffusion models (LDMs). TextMastero introduces two key modules: a glyph conditioning module for fine-grained content control in generating accurate texts, and a latent guidance module for providing comprehensive style information to ensure similarity before and after editing. Both qualitative and quantitative experiments demonstrate that our method surpasses all known existing works in text fidelity and style similarity.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10623
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TextMastero: Mastering High-Quality Scene Text Editing in Diverse Languages and Styles
Wang, Tong
Qu, Xiaochao
Liu, Ting
Computer Vision and Pattern Recognition
Scene text editing aims to modify texts on images while maintaining the style of newly generated text similar to the original. Given an image, a target area, and target text, the task produces an output image with the target text in the selected area, replacing the original. This task has been studied extensively, with initial success using Generative Adversarial Networks (GANs) to balance text fidelity and style similarity. However, GAN-based methods struggled with complex backgrounds or text styles. Recent works leverage diffusion models, showing improved results, yet still face challenges, especially with non-Latin languages like CJK characters (Chinese, Japanese, Korean) that have complex glyphs, often producing inaccurate or unrecognizable characters. To address these issues, we present \emph{TextMastero} - a carefully designed multilingual scene text editing architecture based on latent diffusion models (LDMs). TextMastero introduces two key modules: a glyph conditioning module for fine-grained content control in generating accurate texts, and a latent guidance module for providing comprehensive style information to ensure similarity before and after editing. Both qualitative and quantitative experiments demonstrate that our method surpasses all known existing works in text fidelity and style similarity.
title TextMastero: Mastering High-Quality Scene Text Editing in Diverse Languages and Styles
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2408.10623