IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tang, Yinghao, Liu, Xueding, Zhang, Boyuan, Lan, Tingfeng, Xie, Yupeng, Lao, Jiale, Wang, Yiyao, Li, Haoxuan, Gao, Tingting, Pan, Bo, Weng, Luoxuan, Huang, Xiuqi, Zhu, Minfeng, Feng, Yingchaojie, Luo, Yuyu, Chen, Wei
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908753092149248
author Tang, Yinghao
Liu, Xueding
Zhang, Boyuan
Lan, Tingfeng
Xie, Yupeng
Lao, Jiale
Wang, Yiyao
Li, Haoxuan
Gao, Tingting
Pan, Bo
Weng, Luoxuan
Huang, Xiuqi
Zhu, Minfeng
Feng, Yingchaojie
Luo, Yuyu
Chen, Wei
author_facet Tang, Yinghao
Liu, Xueding
Zhang, Boyuan
Lan, Tingfeng
Xie, Yupeng
Lao, Jiale
Wang, Yiyao
Li, Haoxuan
Gao, Tingting
Pan, Bo
Weng, Luoxuan
Huang, Xiuqi
Zhu, Minfeng
Feng, Yingchaojie
Luo, Yuyu
Chen, Wei
contents Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information. While recent text-to-image (T2I) models can generate aesthetically appealing images, their reliability in generating infographics remains unclear. Generated infographics may appear correct at first glance but contain easily overlooked issues, such as distorted data encoding or incorrect textual content. We present IGENBENCH, the first benchmark for evaluating the reliability of text-to-infographic generation, comprising 600 curated test cases spanning 30 infographic types. We design an automated evaluation framework that decomposes reliability verification into atomic yes/no questions based on a taxonomy of 10 question types. We employ multimodal large language models (MLLMs) to verify each question, yielding question-level accuracy (Q-ACC) and infographic-level accuracy (I-ACC). We comprehensively evaluate 10 state-of-the-art T2I models on IGENBENCH. Our systematic analysis reveals key insights for future model development: (i) a three-tier performance hierarchy with the top model achieving Q-ACC of 0.90 but I-ACC of only 0.49; (ii) data-related dimensions emerging as universal bottlenecks (e.g., Data Completeness: 0.21); and (iii) the challenge of achieving end-to-end correctness across all models. We release IGENBENCH at https://igen-bench.vercel.app/.
format Preprint
id arxiv_https___arxiv_org_abs_2601_04498
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation
Tang, Yinghao
Liu, Xueding
Zhang, Boyuan
Lan, Tingfeng
Xie, Yupeng
Lao, Jiale
Wang, Yiyao
Li, Haoxuan
Gao, Tingting
Pan, Bo
Weng, Luoxuan
Huang, Xiuqi
Zhu, Minfeng
Feng, Yingchaojie
Luo, Yuyu
Chen, Wei
Machine Learning
Computer Vision and Pattern Recognition
Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information. While recent text-to-image (T2I) models can generate aesthetically appealing images, their reliability in generating infographics remains unclear. Generated infographics may appear correct at first glance but contain easily overlooked issues, such as distorted data encoding or incorrect textual content. We present IGENBENCH, the first benchmark for evaluating the reliability of text-to-infographic generation, comprising 600 curated test cases spanning 30 infographic types. We design an automated evaluation framework that decomposes reliability verification into atomic yes/no questions based on a taxonomy of 10 question types. We employ multimodal large language models (MLLMs) to verify each question, yielding question-level accuracy (Q-ACC) and infographic-level accuracy (I-ACC). We comprehensively evaluate 10 state-of-the-art T2I models on IGENBENCH. Our systematic analysis reveals key insights for future model development: (i) a three-tier performance hierarchy with the top model achieving Q-ACC of 0.90 but I-ACC of only 0.49; (ii) data-related dimensions emerging as universal bottlenecks (e.g., Data Completeness: 0.21); and (iii) the challenge of achieving end-to-end correctness across all models. We release IGENBENCH at https://igen-bench.vercel.app/.
title IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.04498