Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Junsung, Lee, Jungbeom, Song, Jongyoon, Yu, Sangwon, Jung, Dahuin, Yoon, Sungroh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910017015250944
author Park, Junsung
Lee, Jungbeom
Song, Jongyoon
Yu, Sangwon
Jung, Dahuin
Yoon, Sungroh
author_facet Park, Junsung
Lee, Jungbeom
Song, Jongyoon
Yu, Sangwon
Jung, Dahuin
Yoon, Sungroh
contents While CLIP has significantly advanced multimodal understanding by bridging vision and language, the inability to grasp negation - such as failing to differentiate concepts like "parking" from "no parking" - poses substantial challenges. By analyzing the data used in the public CLIP model's pre-training, we posit this limitation stems from a lack of negation-inclusive data. To address this, we introduce data generation pipelines that employ a large language model (LLM) and a multimodal LLM to produce negation-inclusive captions. Fine-tuning CLIP with data generated from our pipelines, we develop NegationCLIP, which enhances negation awareness while preserving the generality. Moreover, to enable a comprehensive evaluation of negation understanding, we propose NegRefCOCOg-a benchmark tailored to test VLMs' ability to interpret negation across diverse expressions and positions within a sentence. Experiments on various CLIP architectures validate the effectiveness of our data generation pipelines in enhancing CLIP's ability to perceive negation accurately. Additionally, NegationCLIP's enhanced negation awareness has practical applications across various multimodal tasks, demonstrated by performance gains in text-to-image generation and referring image segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
Park, Junsung
Lee, Jungbeom
Song, Jongyoon
Yu, Sangwon
Jung, Dahuin
Yoon, Sungroh
Computer Vision and Pattern Recognition
Computation and Language
While CLIP has significantly advanced multimodal understanding by bridging vision and language, the inability to grasp negation - such as failing to differentiate concepts like "parking" from "no parking" - poses substantial challenges. By analyzing the data used in the public CLIP model's pre-training, we posit this limitation stems from a lack of negation-inclusive data. To address this, we introduce data generation pipelines that employ a large language model (LLM) and a multimodal LLM to produce negation-inclusive captions. Fine-tuning CLIP with data generated from our pipelines, we develop NegationCLIP, which enhances negation awareness while preserving the generality. Moreover, to enable a comprehensive evaluation of negation understanding, we propose NegRefCOCOg-a benchmark tailored to test VLMs' ability to interpret negation across diverse expressions and positions within a sentence. Experiments on various CLIP architectures validate the effectiveness of our data generation pipelines in enhancing CLIP's ability to perceive negation accurately. Additionally, NegationCLIP's enhanced negation awareness has practical applications across various multimodal tasks, demonstrated by performance gains in text-to-image generation and referring image segmentation.
title Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2501.10913