StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Simonyan, Aleksandr, Jindal, Nipun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909052304359424
author Simonyan, Aleksandr
Jindal, Nipun
author_facet Simonyan, Aleksandr
Jindal, Nipun
contents We present StyleText, a large-scale dataset and benchmark for localized scene-text inpainting with style preservation. StyleText contains 28,518 image-mask-prompt triplets grouped into 9,932 scene families, enabling controlled evaluation of text legibility and visual consistency under shared scene context. We construct the dataset with an automated pipeline that combines LLM prompt templating, Flux-based source generation with key-value (KV) cache injection, OCR-based semantic filtering, polygon mask extraction, and mask-conditioned FluxFill augmentation. We define a reproducible evaluation protocol using normalized OCR metrics (word accuracy and character error rate) and CLIP image-image similarity with explicit preprocessing. A FluxFill+LoRA baseline trained on StyleText improves OCR accuracy substantially over initialization while maintaining scene style consistency, establishing a strong reference point for future comparisons.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17309
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
Simonyan, Aleksandr
Jindal, Nipun
Computer Vision and Pattern Recognition
Artificial Intelligence
We present StyleText, a large-scale dataset and benchmark for localized scene-text inpainting with style preservation. StyleText contains 28,518 image-mask-prompt triplets grouped into 9,932 scene families, enabling controlled evaluation of text legibility and visual consistency under shared scene context. We construct the dataset with an automated pipeline that combines LLM prompt templating, Flux-based source generation with key-value (KV) cache injection, OCR-based semantic filtering, polygon mask extraction, and mask-conditioned FluxFill augmentation. We define a reproducible evaluation protocol using normalized OCR metrics (word accuracy and character error rate) and CLIP image-image similarity with explicit preprocessing. A FluxFill+LoRA baseline trained on StyleText improves OCR accuracy substantially over initialization while maintaining scene style consistency, establishing a strong reference point for future comparisons.
title StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.17309