LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Shitian, Wu, Qilong, Li, Xinyue, Zhang, Bo, Li, Ming, Qin, Qi, Liu, Dongyang, Zhang, Kaipeng, Li, Hongsheng, Qiao, Yu, Gao, Peng, Fu, Bin, Li, Zhen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913762759409664
author Zhao, Shitian
Wu, Qilong
Li, Xinyue
Zhang, Bo
Li, Ming
Qin, Qi
Liu, Dongyang
Zhang, Kaipeng
Li, Hongsheng
Qiao, Yu
Gao, Peng
Fu, Bin
Li, Zhen
author_facet Zhao, Shitian
Wu, Qilong
Li, Xinyue
Zhang, Bo
Li, Ming
Qin, Qi
Liu, Dongyang
Zhang, Kaipeng
Li, Hongsheng
Qiao, Yu
Gao, Peng
Fu, Bin
Li, Zhen
contents We introduce LeX-Art, a comprehensive suite for high-quality text-image synthesis that systematically bridges the gap between prompt expressiveness and text rendering fidelity. Our approach follows a data-centric paradigm, constructing a high-quality data synthesis pipeline based on Deepseek-R1 to curate LeX-10K, a dataset of 10K high-resolution, aesthetically refined 1024$\times$1024 images. Beyond dataset construction, we develop LeX-Enhancer, a robust prompt enrichment model, and train two text-to-image models, LeX-FLUX and LeX-Lumina, achieving state-of-the-art text rendering performance. To systematically evaluate visual text generation, we introduce LeX-Bench, a benchmark that assesses fidelity, aesthetics, and alignment, complemented by Pairwise Normalized Edit Distance (PNED), a novel metric for robust text accuracy evaluation. Experiments demonstrate significant improvements, with LeX-Lumina achieving a 79.81% PNED gain on CreateBench, and LeX-FLUX outperforming baselines in color (+3.18%), positional (+4.45%), and font accuracy (+3.81%). Our codes, models, datasets, and demo are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
Zhao, Shitian
Wu, Qilong
Li, Xinyue
Zhang, Bo
Li, Ming
Qin, Qi
Liu, Dongyang
Zhang, Kaipeng
Li, Hongsheng
Qiao, Yu
Gao, Peng
Fu, Bin
Li, Zhen
Computer Vision and Pattern Recognition
We introduce LeX-Art, a comprehensive suite for high-quality text-image synthesis that systematically bridges the gap between prompt expressiveness and text rendering fidelity. Our approach follows a data-centric paradigm, constructing a high-quality data synthesis pipeline based on Deepseek-R1 to curate LeX-10K, a dataset of 10K high-resolution, aesthetically refined 1024$\times$1024 images. Beyond dataset construction, we develop LeX-Enhancer, a robust prompt enrichment model, and train two text-to-image models, LeX-FLUX and LeX-Lumina, achieving state-of-the-art text rendering performance. To systematically evaluate visual text generation, we introduce LeX-Bench, a benchmark that assesses fidelity, aesthetics, and alignment, complemented by Pairwise Normalized Edit Distance (PNED), a novel metric for robust text accuracy evaluation. Experiments demonstrate significant improvements, with LeX-Lumina achieving a 79.81% PNED gain on CreateBench, and LeX-FLUX outperforming baselines in color (+3.18%), positional (+4.45%), and font accuracy (+3.81%). Our codes, models, datasets, and demo are publicly available.
title LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.21749