GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Tong, Liu, Ting, Qu, Xiaochao, Wu, Chengjing, Liu, Luoqi, Hu, Xiaolin
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909604477140992
author Wang, Tong
Liu, Ting
Qu, Xiaochao
Wu, Chengjing
Liu, Luoqi
Hu, Xiaolin
author_facet Wang, Tong
Liu, Ting
Qu, Xiaochao
Wu, Chengjing
Liu, Luoqi
Hu, Xiaolin
contents Scene text editing, a subfield of image editing, requires modifying texts in images while preserving style consistency and visual coherence with the surrounding environment. While diffusion-based methods have shown promise in text generation, they still struggle to produce high-quality results. These methods often generate distorted or unrecognizable characters, particularly when dealing with complex characters like Chinese. In such systems, characters are composed of intricate stroke patterns and spatial relationships that must be precisely maintained. We present GlyphMastero, a specialized glyph encoder designed to guide the latent diffusion model for generating texts with stroke-level precision. Our key insight is that existing methods, despite using pretrained OCR models for feature extraction, fail to capture the hierarchical nature of text structures - from individual strokes to stroke-level interactions to overall character-level structure. To address this, our glyph encoder explicitly models and captures the cross-level interactions between local-level individual characters and global-level text lines through our novel glyph attention module. Meanwhile, our model implements a feature pyramid network to fuse the multi-scale OCR backbone features at the global-level. Through these cross-level and multi-scale fusions, we obtain more detailed glyph-aware guidance, enabling precise control over the scene text generation process. Our method achieves an 18.02\% improvement in sentence accuracy over the state-of-the-art multi-lingual scene text editing baseline, while simultaneously reducing the text-region Fréchet inception distance by 53.28\%.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04915
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing
Wang, Tong
Liu, Ting
Qu, Xiaochao
Wu, Chengjing
Liu, Luoqi
Hu, Xiaolin
Computer Vision and Pattern Recognition
Scene text editing, a subfield of image editing, requires modifying texts in images while preserving style consistency and visual coherence with the surrounding environment. While diffusion-based methods have shown promise in text generation, they still struggle to produce high-quality results. These methods often generate distorted or unrecognizable characters, particularly when dealing with complex characters like Chinese. In such systems, characters are composed of intricate stroke patterns and spatial relationships that must be precisely maintained. We present GlyphMastero, a specialized glyph encoder designed to guide the latent diffusion model for generating texts with stroke-level precision. Our key insight is that existing methods, despite using pretrained OCR models for feature extraction, fail to capture the hierarchical nature of text structures - from individual strokes to stroke-level interactions to overall character-level structure. To address this, our glyph encoder explicitly models and captures the cross-level interactions between local-level individual characters and global-level text lines through our novel glyph attention module. Meanwhile, our model implements a feature pyramid network to fuse the multi-scale OCR backbone features at the global-level. Through these cross-level and multi-scale fusions, we obtain more detailed glyph-aware guidance, enabling precise control over the scene text generation process. Our method achieves an 18.02\% improvement in sentence accuracy over the state-of-the-art multi-lingual scene text editing baseline, while simultaneously reducing the text-region Fréchet inception distance by 53.28\%.
title GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.04915