Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yufang, Ji, Tao, Sun, Changzhi, Wu, Yuanbin, Zhou, Aimin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912057492766720
author Liu, Yufang
Ji, Tao
Sun, Changzhi
Wu, Yuanbin
Zhou, Aimin
author_facet Liu, Yufang
Ji, Tao
Sun, Changzhi
Wu, Yuanbin
Zhou, Aimin
contents Large Vision-Language Models (LVLMs) have achieved impressive performance, yet research has pointed out a serious issue with object hallucinations within these models. However, there is no clear conclusion as to which part of the model these hallucinations originate from. In this paper, we present an in-depth investigation into the object hallucination problem specifically within the CLIP model, which serves as the backbone for many state-of-the-art vision-language systems. We unveil that even in isolation, the CLIP model is prone to object hallucinations, suggesting that the hallucination problem is not solely due to the interaction between vision and language modalities. To address this, we propose a counterfactual data augmentation method by creating negative samples with a variety of hallucination issues. We demonstrate that our method can effectively mitigate object hallucinations for CLIP model, and we show the the enhanced model can be employed as a visual encoder, effectively alleviating the object hallucination issue in LVLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03176
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
Liu, Yufang
Ji, Tao
Sun, Changzhi
Wu, Yuanbin
Zhou, Aimin
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision-Language Models (LVLMs) have achieved impressive performance, yet research has pointed out a serious issue with object hallucinations within these models. However, there is no clear conclusion as to which part of the model these hallucinations originate from. In this paper, we present an in-depth investigation into the object hallucination problem specifically within the CLIP model, which serves as the backbone for many state-of-the-art vision-language systems. We unveil that even in isolation, the CLIP model is prone to object hallucinations, suggesting that the hallucination problem is not solely due to the interaction between vision and language modalities. To address this, we propose a counterfactual data augmentation method by creating negative samples with a variety of hallucination issues. We demonstrate that our method can effectively mitigate object hallucinations for CLIP model, and we show the the enhanced model can be employed as a visual encoder, effectively alleviating the object hallucination issue in LVLMs.
title Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.03176