Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Jianfei, Zhang, Feng, Sun, Xin, Feng, Chong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908553667674112
author Zhao, Jianfei
Zhang, Feng
Sun, Xin
Feng, Chong
author_facet Zhao, Jianfei
Zhang, Feng
Sun, Xin
Feng, Chong
contents Due to the unidirectional masking mechanism, Decoder-Only models propagate information from left to right. LVLMs (Large Vision-Language Models) follow the same architecture, with visual information gradually integrated into semantic representations during forward propagation. Through systematic analysis, we observe that the majority of the visual information is absorbed into the semantic representations. However, the model's attention distribution does not exhibit sufficient emphasis on semantic representations. This misalignment between the attention distribution and the actual information flow undermines the model's visual understanding ability and contributes to hallucinations. To address this issue, we enhance the model's visual understanding by leveraging the core information embedded in semantic representations. Specifically, we identify attention heads that focus on core semantic representations based on their attention distributions. Then, through a two-stage optimization paradigm, we propagate the advantages of these attention heads across the entire model, aligning the attention distribution with the actual information flow. We evaluate our method on three image captioning benchmarks using five different LVLMs, demonstrating its effectiveness in significantly reducing hallucinations. Further experiments reveal a trade-off between reduced hallucinations and richer details. Notably, our method allows for manual adjustment of the model's conservativeness, enabling flexible control to meet diverse real-world requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14257
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow
Zhao, Jianfei
Zhang, Feng
Sun, Xin
Feng, Chong
Computer Vision and Pattern Recognition
Due to the unidirectional masking mechanism, Decoder-Only models propagate information from left to right. LVLMs (Large Vision-Language Models) follow the same architecture, with visual information gradually integrated into semantic representations during forward propagation. Through systematic analysis, we observe that the majority of the visual information is absorbed into the semantic representations. However, the model's attention distribution does not exhibit sufficient emphasis on semantic representations. This misalignment between the attention distribution and the actual information flow undermines the model's visual understanding ability and contributes to hallucinations. To address this issue, we enhance the model's visual understanding by leveraging the core information embedded in semantic representations. Specifically, we identify attention heads that focus on core semantic representations based on their attention distributions. Then, through a two-stage optimization paradigm, we propagate the advantages of these attention heads across the entire model, aligning the attention distribution with the actual information flow. We evaluate our method on three image captioning benchmarks using five different LVLMs, demonstrating its effectiveness in significantly reducing hallucinations. Further experiments reveal a trade-off between reduced hallucinations and richer details. Notably, our method allows for manual adjustment of the model's conservativeness, enabling flexible control to meet diverse real-world requirements.
title Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.14257