Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Bin, Gao, Dehong, Wang, Yeyuan, Jin, Linbo, Yu, Shanqing, Cai, Xiaoyan, Yang, Libin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915211568480256
author Li, Bin
Gao, Dehong
Wang, Yeyuan
Jin, Linbo
Yu, Shanqing
Cai, Xiaoyan
Yang, Libin
author_facet Li, Bin
Gao, Dehong
Wang, Yeyuan
Jin, Linbo
Yu, Shanqing
Cai, Xiaoyan
Yang, Libin
contents Despite the significant success of Large Vision-Language models(LVLMs), these models still suffer hallucinations when describing images, generating answers that include non-existent objects. It is reported that these models tend to over-focus on certain irrelevant image tokens that do not contain critical information for answering the question and distort the output. To address this, we propose an Instruction-Aligned Visual Attention(IAVA) approach, which identifies irrelevant tokens by comparing changes in attention weights under two different instructions. By applying contrastive decoding, we dynamically adjust the logits generated from original image tokens and irrelevant image tokens, reducing the model's over-attention to irrelevant information. The experimental results demonstrate that IAVA consistently outperforms existing decoding techniques on benchmarks such as MME, POPE, and TextVQA in mitigating object hallucinations. Our IAVA approach is available online at https://github.com/Lee-lab558/IAVA.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18556
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
Li, Bin
Gao, Dehong
Wang, Yeyuan
Jin, Linbo
Yu, Shanqing
Cai, Xiaoyan
Yang, Libin
Computer Vision and Pattern Recognition
Computation and Language
Despite the significant success of Large Vision-Language models(LVLMs), these models still suffer hallucinations when describing images, generating answers that include non-existent objects. It is reported that these models tend to over-focus on certain irrelevant image tokens that do not contain critical information for answering the question and distort the output. To address this, we propose an Instruction-Aligned Visual Attention(IAVA) approach, which identifies irrelevant tokens by comparing changes in attention weights under two different instructions. By applying contrastive decoding, we dynamically adjust the logits generated from original image tokens and irrelevant image tokens, reducing the model's over-attention to irrelevant information. The experimental results demonstrate that IAVA consistently outperforms existing decoding techniques on benchmarks such as MME, POPE, and TextVQA in mitigating object hallucinations. Our IAVA approach is available online at https://github.com/Lee-lab558/IAVA.
title Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.18556