TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Alex Jinpeng, Mao, Dongxing, Zhang, Jiawei, Han, Weiming, Dong, Zhuobai, Li, Linjie, Lin, Yiqi, Yang, Zhengyuan, Qin, Libo, Zhang, Fuwei, Wang, Lijuan, Li, Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915590427377664
author Wang, Alex Jinpeng
Mao, Dongxing
Zhang, Jiawei
Han, Weiming
Dong, Zhuobai
Li, Linjie
Lin, Yiqi
Yang, Zhengyuan
Qin, Libo
Zhang, Fuwei
Wang, Lijuan
Li, Min
author_facet Wang, Alex Jinpeng
Mao, Dongxing
Zhang, Jiawei
Han, Weiming
Dong, Zhuobai
Li, Linjie
Lin, Yiqi
Yang, Zhengyuan
Qin, Libo
Zhang, Fuwei
Wang, Lijuan
Li, Min
contents Text-conditioned image generation has gained significant attention in recent years and are processing increasingly longer and comprehensive text prompt. In everyday life, dense and intricate text appears in contexts like advertisements, infographics, and signage, where the integration of both text and visuals is essential for conveying complex information. However, despite these advances, the generation of images containing long-form text remains a persistent challenge, largely due to the limitations of existing datasets, which often focus on shorter and simpler text. To address this gap, we introduce TextAtlas5M, a novel dataset specifically designed to evaluate long-text rendering in text-conditioned image generation. Our dataset consists of 5 million long-text generated and collected images across diverse data types, enabling comprehensive evaluation of large-scale generative models on long-text image generation. We further curate 3000 human-improved test set TextAtlasEval across 3 data domains, establishing one of the most extensive benchmarks for text-conditioned generation. Evaluations suggest that the TextAtlasEval benchmarks present significant challenges even for the most advanced proprietary models (e.g. GPT4o with DallE-3), while their open-source counterparts show an even larger performance gap. These evidences position TextAtlas5M as a valuable dataset for training and evaluating future-generation text-conditioned image generation models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07870
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
Wang, Alex Jinpeng
Mao, Dongxing
Zhang, Jiawei
Han, Weiming
Dong, Zhuobai
Li, Linjie
Lin, Yiqi
Yang, Zhengyuan
Qin, Libo
Zhang, Fuwei
Wang, Lijuan
Li, Min
Computer Vision and Pattern Recognition
Text-conditioned image generation has gained significant attention in recent years and are processing increasingly longer and comprehensive text prompt. In everyday life, dense and intricate text appears in contexts like advertisements, infographics, and signage, where the integration of both text and visuals is essential for conveying complex information. However, despite these advances, the generation of images containing long-form text remains a persistent challenge, largely due to the limitations of existing datasets, which often focus on shorter and simpler text. To address this gap, we introduce TextAtlas5M, a novel dataset specifically designed to evaluate long-text rendering in text-conditioned image generation. Our dataset consists of 5 million long-text generated and collected images across diverse data types, enabling comprehensive evaluation of large-scale generative models on long-text image generation. We further curate 3000 human-improved test set TextAtlasEval across 3 data domains, establishing one of the most extensive benchmarks for text-conditioned generation. Evaluations suggest that the TextAtlasEval benchmarks present significant challenges even for the most advanced proprietary models (e.g. GPT4o with DallE-3), while their open-source counterparts show an even larger performance gap. These evidences position TextAtlas5M as a valuable dataset for training and evaluating future-generation text-conditioned image generation models.
title TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.07870