KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Hsin-Ping, Wang, Xinyi, Bitton, Yonatan, Taitelbaum, Hagai, Tomar, Gaurav Singh, Chang, Ming-Wei, Jia, Xuhui, Chan, Kelvin C. K., Hu, Hexiang, Su, Yu-Chuan, Yang, Ming-Hsuan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908468863041536
author Huang, Hsin-Ping
Wang, Xinyi
Bitton, Yonatan
Taitelbaum, Hagai
Tomar, Gaurav Singh
Chang, Ming-Wei
Jia, Xuhui
Chan, Kelvin C. K.
Hu, Hexiang
Su, Yu-Chuan
Yang, Ming-Hsuan
author_facet Huang, Hsin-Ping
Wang, Xinyi
Bitton, Yonatan
Taitelbaum, Hagai
Tomar, Gaurav Singh
Chang, Ming-Wei
Jia, Xuhui
Chan, Kelvin C. K.
Hu, Hexiang
Su, Yu-Chuan
Yang, Ming-Hsuan
contents Recent advances in text-to-image generation have improved the quality of synthesized images, but evaluations mainly focus on aesthetics or alignment with text prompts. Thus, it remains unclear whether these models can accurately represent a wide variety of realistic visual entities. To bridge this gap, we propose KITTEN, a benchmark for Knowledge-InTensive image generaTion on real-world ENtities. Using KITTEN, we conduct a systematic study of the latest text-to-image models and retrieval-augmented models, focusing on their ability to generate real-world visual entities, such as landmarks and animals. Analysis using carefully designed human evaluations, automatic metrics, and MLLM evaluations show that even advanced text-to-image models fail to generate accurate visual details of entities. While retrieval-augmented models improve entity fidelity by incorporating reference images, they tend to over-rely on them and struggle to create novel configurations of the entity in creative text prompts.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11824
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities
Huang, Hsin-Ping
Wang, Xinyi
Bitton, Yonatan
Taitelbaum, Hagai
Tomar, Gaurav Singh
Chang, Ming-Wei
Jia, Xuhui
Chan, Kelvin C. K.
Hu, Hexiang
Su, Yu-Chuan
Yang, Ming-Hsuan
Computer Vision and Pattern Recognition
Recent advances in text-to-image generation have improved the quality of synthesized images, but evaluations mainly focus on aesthetics or alignment with text prompts. Thus, it remains unclear whether these models can accurately represent a wide variety of realistic visual entities. To bridge this gap, we propose KITTEN, a benchmark for Knowledge-InTensive image generaTion on real-world ENtities. Using KITTEN, we conduct a systematic study of the latest text-to-image models and retrieval-augmented models, focusing on their ability to generate real-world visual entities, such as landmarks and animals. Analysis using carefully designed human evaluations, automatic metrics, and MLLM evaluations show that even advanced text-to-image models fail to generate accurate visual details of entities. While retrieval-augmented models improve entity fidelity by incorporating reference images, they tend to over-rely on them and struggle to create novel configurations of the entity in creative text prompts.
title KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.11824