Attention to details, logits to truth: visual-aware attention and logits enhancement to mitigate hallucinations in LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jingyi, Li, Fei, Liu, Rujie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908825578110976
author Wang, Jingyi
Li, Fei
Liu, Rujie
author_facet Wang, Jingyi
Li, Fei
Liu, Rujie
contents Existing Large Vision-Language Models (LVLMs) exhibit insufficient visual attention, leading to hallucinations. To alleviate this problem, some previous studies adjust and amplify visual attention. These methods present a limitation that boosting attention for all visual tokens inevitably increases attention to task irrelevant tokens. To tackle this challenge, we propose a training free attentional intervention algorithm to enhance the attention of task-relevant tokens based on the argument that task-relevant tokens generally demonstrate high visual-textual similarities. Specifically, the vision-text cross-attention submatrices, which represent visual-textual correlations, are extracted to construct the reweighting matrices to reallocate attention. Besides, to enhance the contribution of visual tokens, we inject visual attention values into the beam search decoding to identify solutions with higher visual attention. Extensive experiments demonstrate that this method significantly reduces hallucinations across mainstream LVLMs, while preserving the accuracy and coherence of generated content.
format Preprint
id arxiv_https___arxiv_org_abs_2602_09521
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Attention to details, logits to truth: visual-aware attention and logits enhancement to mitigate hallucinations in LVLMs
Wang, Jingyi
Li, Fei
Liu, Rujie
Computer Vision and Pattern Recognition
Existing Large Vision-Language Models (LVLMs) exhibit insufficient visual attention, leading to hallucinations. To alleviate this problem, some previous studies adjust and amplify visual attention. These methods present a limitation that boosting attention for all visual tokens inevitably increases attention to task irrelevant tokens. To tackle this challenge, we propose a training free attentional intervention algorithm to enhance the attention of task-relevant tokens based on the argument that task-relevant tokens generally demonstrate high visual-textual similarities. Specifically, the vision-text cross-attention submatrices, which represent visual-textual correlations, are extracted to construct the reweighting matrices to reallocate attention. Besides, to enhance the contribution of visual tokens, we inject visual attention values into the beam search decoding to identify solutions with higher visual attention. Extensive experiments demonstrate that this method significantly reduces hallucinations across mainstream LVLMs, while preserving the accuracy and coherence of generated content.
title Attention to details, logits to truth: visual-aware attention and logits enhancement to mitigate hallucinations in LVLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.09521