A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Udandarao, Vishaal, Cherti, Mehdi, Karthik, Shyamgopal, Jitsev, Jenia, Albanie, Samuel, Bethge, Matthias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912422322765824
author Udandarao, Vishaal
Cherti, Mehdi
Karthik, Shyamgopal
Jitsev, Jenia
Albanie, Samuel
Bethge, Matthias
author_facet Udandarao, Vishaal
Cherti, Mehdi
Karthik, Shyamgopal
Jitsev, Jenia
Albanie, Samuel
Bethge, Matthias
contents We investigate 17 benchmarks (e.g. SugarCREPE, VALSE) commonly used for measuring compositional understanding capabilities of vision-language models (VLMs). We scrutinize design choices in their construction, including data source (e.g. MS-COCO) and curation procedures (e.g. constructing negative images/captions), uncovering several inherent biases across most benchmarks. We find that blind heuristics (e.g. token-length, log-likelihood under a language model) perform on par with CLIP models, indicating that these benchmarks do not effectively measure compositional understanding. We demonstrate that the underlying factor is a distribution asymmetry between positive and negative images/captions, induced by the benchmark construction procedures. To mitigate these issues, we provide a few key recommendations for constructing more robust vision-language compositional understanding benchmarks, that would be less prone to such simple attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08227
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
Udandarao, Vishaal
Cherti, Mehdi
Karthik, Shyamgopal
Jitsev, Jenia
Albanie, Samuel
Bethge, Matthias
Computer Vision and Pattern Recognition
We investigate 17 benchmarks (e.g. SugarCREPE, VALSE) commonly used for measuring compositional understanding capabilities of vision-language models (VLMs). We scrutinize design choices in their construction, including data source (e.g. MS-COCO) and curation procedures (e.g. constructing negative images/captions), uncovering several inherent biases across most benchmarks. We find that blind heuristics (e.g. token-length, log-likelihood under a language model) perform on par with CLIP models, indicating that these benchmarks do not effectively measure compositional understanding. We demonstrate that the underlying factor is a distribution asymmetry between positive and negative images/captions, induced by the benchmark construction procedures. To mitigate these issues, we provide a few key recommendations for constructing more robust vision-language compositional understanding benchmarks, that would be less prone to such simple attacks.
title A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.08227