A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912422322765824 |
|---|---|
| author | Udandarao, Vishaal Cherti, Mehdi Karthik, Shyamgopal Jitsev, Jenia Albanie, Samuel Bethge, Matthias |
| author_facet | Udandarao, Vishaal Cherti, Mehdi Karthik, Shyamgopal Jitsev, Jenia Albanie, Samuel Bethge, Matthias |
| contents | We investigate 17 benchmarks (e.g. SugarCREPE, VALSE) commonly used for measuring compositional understanding capabilities of vision-language models (VLMs). We scrutinize design choices in their construction, including data source (e.g. MS-COCO) and curation procedures (e.g. constructing negative images/captions), uncovering several inherent biases across most benchmarks. We find that blind heuristics (e.g. token-length, log-likelihood under a language model) perform on par with CLIP models, indicating that these benchmarks do not effectively measure compositional understanding. We demonstrate that the underlying factor is a distribution asymmetry between positive and negative images/captions, induced by the benchmark construction procedures. To mitigate these issues, we provide a few key recommendations for constructing more robust vision-language compositional understanding benchmarks, that would be less prone to such simple attacks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_08227 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Udandarao, Vishaal Cherti, Mehdi Karthik, Shyamgopal Jitsev, Jenia Albanie, Samuel Bethge, Matthias Computer Vision and Pattern Recognition We investigate 17 benchmarks (e.g. SugarCREPE, VALSE) commonly used for measuring compositional understanding capabilities of vision-language models (VLMs). We scrutinize design choices in their construction, including data source (e.g. MS-COCO) and curation procedures (e.g. constructing negative images/captions), uncovering several inherent biases across most benchmarks. We find that blind heuristics (e.g. token-length, log-likelihood under a language model) perform on par with CLIP models, indicating that these benchmarks do not effectively measure compositional understanding. We demonstrate that the underlying factor is a distribution asymmetry between positive and negative images/captions, induced by the benchmark construction procedures. To mitigate these issues, we provide a few key recommendations for constructing more robust vision-language compositional understanding benchmarks, that would be less prone to such simple attacks. |
| title | A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.08227 |