T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Ziwei, He, Wanggui, Long, Quanyu, Wang, Yandi, Li, Haoyuan, Yu, Zhelun, Shu, Fangxun, Chan, Long, Jiang, Hao, Wu, Fei, Gan, Leilei
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911599383543808
author Huang, Ziwei
He, Wanggui
Long, Quanyu
Wang, Yandi
Li, Haoyuan
Yu, Zhelun
Shu, Fangxun
Chan, Long
Jiang, Hao
Wu, Fei
Gan, Leilei
author_facet Huang, Ziwei
He, Wanggui
Long, Quanyu
Wang, Yandi
Li, Haoyuan
Yu, Zhelun
Shu, Fangxun
Chan, Long
Jiang, Hao
Wu, Fei
Gan, Leilei
contents Evaluating the quality of synthesized images remains a significant challenge in the development of text-to-image (T2I) generation. Most existing studies in this area primarily focus on evaluating text-image alignment, image quality, and object composition capabilities, with comparatively fewer studies addressing the evaluation of the factuality of T2I models, particularly when the concepts involved are knowledge-intensive. To mitigate this gap, we present T2I-FactualBench in this work - the largest benchmark to date in terms of the number of concepts and prompts specifically designed to evaluate the factuality of knowledge-intensive concept generation. T2I-FactualBench consists of a three-tiered knowledge-intensive text-to-image generation framework, ranging from the basic memorization of individual knowledge concepts to the more complex composition of multiple knowledge concepts. We further introduce a multi-round visual question answering (VQA) based evaluation framework to assess the factuality of three-tiered knowledge-intensive text-to-image generation tasks. Experiments on T2I-FactualBench indicate that current state-of-the-art (SOTA) T2I models still leave significant room for improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04300
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
Huang, Ziwei
He, Wanggui
Long, Quanyu
Wang, Yandi
Li, Haoyuan
Yu, Zhelun
Shu, Fangxun
Chan, Long
Jiang, Hao
Wu, Fei
Gan, Leilei
Computer Vision and Pattern Recognition
Artificial Intelligence
Evaluating the quality of synthesized images remains a significant challenge in the development of text-to-image (T2I) generation. Most existing studies in this area primarily focus on evaluating text-image alignment, image quality, and object composition capabilities, with comparatively fewer studies addressing the evaluation of the factuality of T2I models, particularly when the concepts involved are knowledge-intensive. To mitigate this gap, we present T2I-FactualBench in this work - the largest benchmark to date in terms of the number of concepts and prompts specifically designed to evaluate the factuality of knowledge-intensive concept generation. T2I-FactualBench consists of a three-tiered knowledge-intensive text-to-image generation framework, ranging from the basic memorization of individual knowledge concepts to the more complex composition of multiple knowledge concepts. We further introduce a multi-round visual question answering (VQA) based evaluation framework to assess the factuality of three-tiered knowledge-intensive text-to-image generation tasks. Experiments on T2I-FactualBench indicate that current state-of-the-art (SOTA) T2I models still leave significant room for improvement.
title T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.04300