Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: An, Wenbin, Tian, Feng, Leng, Sicong, Nie, Jiahao, Lin, Haonan, Wang, QianYing, Chen, Ping, Zhang, Xiaoqin, Lu, Shijian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913735223803904
author An, Wenbin
Tian, Feng
Leng, Sicong
Nie, Jiahao
Lin, Haonan
Wang, QianYing
Chen, Ping
Zhang, Xiaoqin
Lu, Shijian
author_facet An, Wenbin
Tian, Feng
Leng, Sicong
Nie, Jiahao
Lin, Haonan
Wang, QianYing
Chen, Ping
Zhang, Xiaoqin
Lu, Shijian
contents Despite great success across various multimodal tasks, Large Vision-Language Models (LVLMs) often encounter object hallucinations with generated textual responses being inconsistent with the actual objects in images. We examine different LVLMs and pinpoint that one root cause of object hallucinations lies with deficient attention on discriminative image features. Specifically, LVLMs often predominantly attend to prompt-irrelevant global features instead of prompt-relevant local features, undermining their visual grounding capacity and leading to object hallucinations. We propose Assembly of Global and Local Attention (AGLA), a training-free and plug-and-play approach that mitigates hallucinations by assembling global features for response generation and local features for visual discrimination simultaneously. Specifically, we introduce an image-prompt matching scheme that captures prompt-relevant local features from images, leading to an augmented view of the input image where prompt-relevant content is highlighted while irrelevant distractions are suppressed. Hallucinations can thus be mitigated with a calibrated logit distribution that is from generative global features of the original image and discriminative local features of the augmented image. Extensive experiments show the superiority of AGLA in LVLM hallucination mitigation, demonstrating its wide applicability across both discriminative and generative tasks. Our code is available at https://github.com/Lackel/AGLA.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12718
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
An, Wenbin
Tian, Feng
Leng, Sicong
Nie, Jiahao
Lin, Haonan
Wang, QianYing
Chen, Ping
Zhang, Xiaoqin
Lu, Shijian
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Despite great success across various multimodal tasks, Large Vision-Language Models (LVLMs) often encounter object hallucinations with generated textual responses being inconsistent with the actual objects in images. We examine different LVLMs and pinpoint that one root cause of object hallucinations lies with deficient attention on discriminative image features. Specifically, LVLMs often predominantly attend to prompt-irrelevant global features instead of prompt-relevant local features, undermining their visual grounding capacity and leading to object hallucinations. We propose Assembly of Global and Local Attention (AGLA), a training-free and plug-and-play approach that mitigates hallucinations by assembling global features for response generation and local features for visual discrimination simultaneously. Specifically, we introduce an image-prompt matching scheme that captures prompt-relevant local features from images, leading to an augmented view of the input image where prompt-relevant content is highlighted while irrelevant distractions are suppressed. Hallucinations can thus be mitigated with a calibrated logit distribution that is from generative global features of the original image and discriminative local features of the augmented image. Extensive experiments show the superiority of AGLA in LVLM hallucination mitigation, demonstrating its wide applicability across both discriminative and generative tasks. Our code is available at https://github.com/Lackel/AGLA.
title Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local Attention
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2406.12718