NegVQA: Can Vision Language Models Understand Negation?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuhui, Su, Yuchang, Liu, Yiming, Yeung-Levy, Serena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912400905601024
author Zhang, Yuhui
Su, Yuchang
Liu, Yiming
Yeung-Levy, Serena
author_facet Zhang, Yuhui
Su, Yuchang
Liu, Yiming
Yeung-Levy, Serena
contents Negation is a fundamental linguistic phenomenon that can entirely reverse the meaning of a sentence. As vision language models (VLMs) continue to advance and are deployed in high-stakes applications, assessing their ability to comprehend negation becomes essential. To address this, we introduce NegVQA, a visual question answering (VQA) benchmark consisting of 7,379 two-choice questions covering diverse negation scenarios and image-question distributions. We construct NegVQA by leveraging large language models to generate negated versions of questions from existing VQA datasets. Evaluating 20 state-of-the-art VLMs across seven model families, we find that these models struggle significantly with negation, exhibiting a substantial performance drop compared to their responses to the original questions. Furthermore, we uncover a U-shaped scaling trend, where increasing model size initially degrades performance on NegVQA before leading to improvements. Our benchmark reveals critical gaps in VLMs' negation understanding and offers insights into future VLM development. Project page available at https://yuhui-zh15.github.io/NegVQA/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22946
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle NegVQA: Can Vision Language Models Understand Negation?
Zhang, Yuhui
Su, Yuchang
Liu, Yiming
Yeung-Levy, Serena
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Computers and Society
Machine Learning
Negation is a fundamental linguistic phenomenon that can entirely reverse the meaning of a sentence. As vision language models (VLMs) continue to advance and are deployed in high-stakes applications, assessing their ability to comprehend negation becomes essential. To address this, we introduce NegVQA, a visual question answering (VQA) benchmark consisting of 7,379 two-choice questions covering diverse negation scenarios and image-question distributions. We construct NegVQA by leveraging large language models to generate negated versions of questions from existing VQA datasets. Evaluating 20 state-of-the-art VLMs across seven model families, we find that these models struggle significantly with negation, exhibiting a substantial performance drop compared to their responses to the original questions. Furthermore, we uncover a U-shaped scaling trend, where increasing model size initially degrades performance on NegVQA before leading to improvements. Our benchmark reveals critical gaps in VLMs' negation understanding and offers insights into future VLM development. Project page available at https://yuhui-zh15.github.io/NegVQA/.
title NegVQA: Can Vision Language Models Understand Negation?
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Computers and Society
Machine Learning
url https://arxiv.org/abs/2505.22946