Evaluating Compositional Scene Understanding in Multimodal Generative Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fu, Shuhao, Lee, Andrew Jun, Wang, Anna, Momennejad, Ida, Bihl, Trevor, Lu, Hongjing, Webb, Taylor W.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908289870069760
author Fu, Shuhao
Lee, Andrew Jun
Wang, Anna
Momennejad, Ida
Bihl, Trevor
Lu, Hongjing
Webb, Taylor W.
author_facet Fu, Shuhao
Lee, Andrew Jun
Wang, Anna
Momennejad, Ida
Bihl, Trevor
Lu, Hongjing
Webb, Taylor W.
contents The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential for computer vision systems to reflect and exploit this compositionality to achieve robust and generalizable scene understanding. While major strides have been made toward the development of general-purpose, multimodal generative models, including both text-to-image models and multimodal vision-language models, it remains unclear whether these systems are capable of accurately generating and interpreting scenes involving the composition of multiple objects and relations. In this work, we present an evaluation of the compositional visual processing capabilities in the current generation of text-to-image (DALL-E 3) and multimodal vision-language models (GPT-4V, GPT-4o, Claude Sonnet 3.5, QWEN2-VL-72B, and InternVL2.5-38B), and compare the performance of these systems to human participants. The results suggest that these systems display some ability to solve compositional and relational tasks, showing notable improvements over the previous generation of multimodal models, but with performance nevertheless well below the level of human participants, particularly for more complex scenes involving many ($>5$) objects and multiple relations. These results highlight the need for further progress toward compositional understanding of visual scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23125
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Compositional Scene Understanding in Multimodal Generative Models
Fu, Shuhao
Lee, Andrew Jun
Wang, Anna
Momennejad, Ida
Bihl, Trevor
Lu, Hongjing
Webb, Taylor W.
Computer Vision and Pattern Recognition
Artificial Intelligence
The visual world is fundamentally compositional. Visual scenes are defined by the composition of objects and their relations. Hence, it is essential for computer vision systems to reflect and exploit this compositionality to achieve robust and generalizable scene understanding. While major strides have been made toward the development of general-purpose, multimodal generative models, including both text-to-image models and multimodal vision-language models, it remains unclear whether these systems are capable of accurately generating and interpreting scenes involving the composition of multiple objects and relations. In this work, we present an evaluation of the compositional visual processing capabilities in the current generation of text-to-image (DALL-E 3) and multimodal vision-language models (GPT-4V, GPT-4o, Claude Sonnet 3.5, QWEN2-VL-72B, and InternVL2.5-38B), and compare the performance of these systems to human participants. The results suggest that these systems display some ability to solve compositional and relational tasks, showing notable improvements over the previous generation of multimodal models, but with performance nevertheless well below the level of human participants, particularly for more complex scenes involving many ($>5$) objects and multiple relations. These results highlight the need for further progress toward compositional understanding of visual scenes.
title Evaluating Compositional Scene Understanding in Multimodal Generative Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.23125