CIVET: Systematic Evaluation of Understanding in VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rizzoli, Massimo, Alghisi, Simone, Khomyn, Olha, Roccabruna, Gabriel, Mousavi, Seyed Mahed, Riccardi, Giuseppe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912439350591488
author Rizzoli, Massimo
Alghisi, Simone
Khomyn, Olha
Roccabruna, Gabriel
Mousavi, Seyed Mahed
Riccardi, Giuseppe
author_facet Rizzoli, Massimo
Alghisi, Simone
Khomyn, Olha
Roccabruna, Gabriel
Mousavi, Seyed Mahed
Riccardi, Giuseppe
contents While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains understudied. To investigate the understanding of VLMs, we study their capability regarding object properties and relations in a controlled and interpretable manner. To this scope, we introduce CIVET, a novel and extensible framework for systematiC evaluatIon Via controllEd sTimuli. CIVET addresses the lack of standardized systematic evaluation for assessing VLMs' understanding, enabling researchers to test hypotheses with statistical rigor. With CIVET, we evaluate five state-of-the-art VLMs on exhaustive sets of stimuli, free from annotation noise, dataset-specific biases, and uncontrolled scene complexity. Our findings reveal that 1) current VLMs can accurately recognize only a limited set of basic object properties; 2) their performance heavily depends on the position of the object in the scene; 3) they struggle to understand basic relations among objects. Furthermore, a comparative evaluation with human annotators reveals that VLMs still fall short of achieving human-level accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05146
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CIVET: Systematic Evaluation of Understanding in VLMs
Rizzoli, Massimo
Alghisi, Simone
Khomyn, Olha
Roccabruna, Gabriel
Mousavi, Seyed Mahed
Riccardi, Giuseppe
Computer Vision and Pattern Recognition
Computation and Language
While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains understudied. To investigate the understanding of VLMs, we study their capability regarding object properties and relations in a controlled and interpretable manner. To this scope, we introduce CIVET, a novel and extensible framework for systematiC evaluatIon Via controllEd sTimuli. CIVET addresses the lack of standardized systematic evaluation for assessing VLMs' understanding, enabling researchers to test hypotheses with statistical rigor. With CIVET, we evaluate five state-of-the-art VLMs on exhaustive sets of stimuli, free from annotation noise, dataset-specific biases, and uncontrolled scene complexity. Our findings reveal that 1) current VLMs can accurately recognize only a limited set of basic object properties; 2) their performance heavily depends on the position of the object in the scene; 3) they struggle to understand basic relations among objects. Furthermore, a comparative evaluation with human annotators reveals that VLMs still fall short of achieving human-level accuracy.
title CIVET: Systematic Evaluation of Understanding in VLMs
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2506.05146