TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yu, Zhang, Jielei, Chen, Pengyu, Wang, Weihang, Gao, Longwen, Li, Peiyi, Qiao, Qian, Lian, Zhouhui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910049890205696
author Xie, Yu
Zhang, Jielei
Chen, Pengyu
Wang, Weihang
Gao, Longwen
Li, Peiyi
Qiao, Qian
Lian, Zhouhui
author_facet Xie, Yu
Zhang, Jielei
Chen, Pengyu
Wang, Weihang
Gao, Longwen
Li, Peiyi
Qiao, Qian
Lian, Zhouhui
contents Diffusion-based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large-scale annotated data to support multilingual generation. In this work, we revisit the necessity of complex auxiliary modules and further explore an approach that simultaneously ensures glyph accuracy and achieves high-fidelity scene integration, by leveraging diffusion models' inherent capabilities for contextual reasoning. To this end, we introduce TextFlux, a DiT-based framework that enables multilingual scene text synthesis. The advantages of TextFlux can be summarized as follows: (1) OCR-free model architecture. TextFlux eliminates the need for OCR encoders (additional visual conditioning modules) that are specifically used to extract visual text-related features. (2) Strong multilingual scalability. TextFlux is effective in low-resource multilingual settings, and achieves strong performance in newly added languages with fewer than 1,000 samples. (3) Streamlined training setup. TextFlux is trained with only 1% of the training data required by competing methods. (4) Controllable multi-line text generation. TextFlux offers flexible multi-line synthesis with precise line-level control, outperforming methods restricted to single-line or rigid layouts. Extensive experiments and visualizations demonstrate that TextFlux outperforms previous methods in both qualitative and quantitative evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17778
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
Xie, Yu
Zhang, Jielei
Chen, Pengyu
Wang, Weihang
Gao, Longwen
Li, Peiyi
Qiao, Qian
Lian, Zhouhui
Computer Vision and Pattern Recognition
Diffusion-based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large-scale annotated data to support multilingual generation. In this work, we revisit the necessity of complex auxiliary modules and further explore an approach that simultaneously ensures glyph accuracy and achieves high-fidelity scene integration, by leveraging diffusion models' inherent capabilities for contextual reasoning. To this end, we introduce TextFlux, a DiT-based framework that enables multilingual scene text synthesis. The advantages of TextFlux can be summarized as follows: (1) OCR-free model architecture. TextFlux eliminates the need for OCR encoders (additional visual conditioning modules) that are specifically used to extract visual text-related features. (2) Strong multilingual scalability. TextFlux is effective in low-resource multilingual settings, and achieves strong performance in newly added languages with fewer than 1,000 samples. (3) Streamlined training setup. TextFlux is trained with only 1% of the training data required by competing methods. (4) Controllable multi-line text generation. TextFlux offers flexible multi-line synthesis with precise line-level control, outperforming methods restricted to single-line or rigid layouts. Extensive experiments and visualizations demonstrate that TextFlux outperforms previous methods in both qualitative and quantitative evaluations.
title TextFlux: An OCR-Free DiT Model for High-Fidelity Multilingual Scene Text Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17778