KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908468863041536 |
|---|---|
| author | Huang, Hsin-Ping Wang, Xinyi Bitton, Yonatan Taitelbaum, Hagai Tomar, Gaurav Singh Chang, Ming-Wei Jia, Xuhui Chan, Kelvin C. K. Hu, Hexiang Su, Yu-Chuan Yang, Ming-Hsuan |
| author_facet | Huang, Hsin-Ping Wang, Xinyi Bitton, Yonatan Taitelbaum, Hagai Tomar, Gaurav Singh Chang, Ming-Wei Jia, Xuhui Chan, Kelvin C. K. Hu, Hexiang Su, Yu-Chuan Yang, Ming-Hsuan |
| contents | Recent advances in text-to-image generation have improved the quality of synthesized images, but evaluations mainly focus on aesthetics or alignment with text prompts. Thus, it remains unclear whether these models can accurately represent a wide variety of realistic visual entities. To bridge this gap, we propose KITTEN, a benchmark for Knowledge-InTensive image generaTion on real-world ENtities. Using KITTEN, we conduct a systematic study of the latest text-to-image models and retrieval-augmented models, focusing on their ability to generate real-world visual entities, such as landmarks and animals. Analysis using carefully designed human evaluations, automatic metrics, and MLLM evaluations show that even advanced text-to-image models fail to generate accurate visual details of entities. While retrieval-augmented models improve entity fidelity by incorporating reference images, they tend to over-rely on them and struggle to create novel configurations of the entity in creative text prompts. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_11824 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities Huang, Hsin-Ping Wang, Xinyi Bitton, Yonatan Taitelbaum, Hagai Tomar, Gaurav Singh Chang, Ming-Wei Jia, Xuhui Chan, Kelvin C. K. Hu, Hexiang Su, Yu-Chuan Yang, Ming-Hsuan Computer Vision and Pattern Recognition Recent advances in text-to-image generation have improved the quality of synthesized images, but evaluations mainly focus on aesthetics or alignment with text prompts. Thus, it remains unclear whether these models can accurately represent a wide variety of realistic visual entities. To bridge this gap, we propose KITTEN, a benchmark for Knowledge-InTensive image generaTion on real-world ENtities. Using KITTEN, we conduct a systematic study of the latest text-to-image models and retrieval-augmented models, focusing on their ability to generate real-world visual entities, such as landmarks and animals. Analysis using carefully designed human evaluations, automatic metrics, and MLLM evaluations show that even advanced text-to-image models fail to generate accurate visual details of entities. While retrieval-augmented models improve entity fidelity by incorporating reference images, they tend to over-rely on them and struggle to create novel configurations of the entity in creative text prompts. |
| title | KITTEN: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.11824 |