LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913762759409664 |
|---|---|
| author | Zhao, Shitian Wu, Qilong Li, Xinyue Zhang, Bo Li, Ming Qin, Qi Liu, Dongyang Zhang, Kaipeng Li, Hongsheng Qiao, Yu Gao, Peng Fu, Bin Li, Zhen |
| author_facet | Zhao, Shitian Wu, Qilong Li, Xinyue Zhang, Bo Li, Ming Qin, Qi Liu, Dongyang Zhang, Kaipeng Li, Hongsheng Qiao, Yu Gao, Peng Fu, Bin Li, Zhen |
| contents | We introduce LeX-Art, a comprehensive suite for high-quality text-image synthesis that systematically bridges the gap between prompt expressiveness and text rendering fidelity. Our approach follows a data-centric paradigm, constructing a high-quality data synthesis pipeline based on Deepseek-R1 to curate LeX-10K, a dataset of 10K high-resolution, aesthetically refined 1024$\times$1024 images. Beyond dataset construction, we develop LeX-Enhancer, a robust prompt enrichment model, and train two text-to-image models, LeX-FLUX and LeX-Lumina, achieving state-of-the-art text rendering performance. To systematically evaluate visual text generation, we introduce LeX-Bench, a benchmark that assesses fidelity, aesthetics, and alignment, complemented by Pairwise Normalized Edit Distance (PNED), a novel metric for robust text accuracy evaluation. Experiments demonstrate significant improvements, with LeX-Lumina achieving a 79.81% PNED gain on CreateBench, and LeX-FLUX outperforming baselines in color (+3.18%), positional (+4.45%), and font accuracy (+3.81%). Our codes, models, datasets, and demo are publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_21749 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis Zhao, Shitian Wu, Qilong Li, Xinyue Zhang, Bo Li, Ming Qin, Qi Liu, Dongyang Zhang, Kaipeng Li, Hongsheng Qiao, Yu Gao, Peng Fu, Bin Li, Zhen Computer Vision and Pattern Recognition We introduce LeX-Art, a comprehensive suite for high-quality text-image synthesis that systematically bridges the gap between prompt expressiveness and text rendering fidelity. Our approach follows a data-centric paradigm, constructing a high-quality data synthesis pipeline based on Deepseek-R1 to curate LeX-10K, a dataset of 10K high-resolution, aesthetically refined 1024$\times$1024 images. Beyond dataset construction, we develop LeX-Enhancer, a robust prompt enrichment model, and train two text-to-image models, LeX-FLUX and LeX-Lumina, achieving state-of-the-art text rendering performance. To systematically evaluate visual text generation, we introduce LeX-Bench, a benchmark that assesses fidelity, aesthetics, and alignment, complemented by Pairwise Normalized Edit Distance (PNED), a novel metric for robust text accuracy evaluation. Experiments demonstrate significant improvements, with LeX-Lumina achieving a 79.81% PNED gain on CreateBench, and LeX-FLUX outperforming baselines in color (+3.18%), positional (+4.45%), and font accuracy (+3.81%). Our codes, models, datasets, and demo are publicly available. |
| title | LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.21749 |