V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Nan, Zhang, Zhenyu, Lin, Xixun, Wang, Kun, Shang, Yanmin, Gu, Naibin, Wang, Shuohuan, Sun, Yu, Wu, Hua, Wang, Haifeng, Cao, Yanan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914178606825472
author Sun, Nan
Zhang, Zhenyu
Lin, Xixun
Wang, Kun
Shang, Yanmin
Gu, Naibin
Wang, Shuohuan
Sun, Yu
Wu, Hua
Wang, Haifeng
Cao, Yanan
author_facet Sun, Nan
Zhang, Zhenyu
Lin, Xixun
Wang, Kun
Shang, Yanmin
Gu, Naibin
Wang, Shuohuan
Sun, Yu
Wu, Hua
Wang, Haifeng
Cao, Yanan
contents Multimodal Large Language Models (MLLMs) excel in numerous vision-language tasks yet suffer from hallucinations, producing content inconsistent with input visuals, that undermine reliability in precision-sensitive domains. This issue stems from a fundamental problem of visual neglect, where models fail to adequately prioritize input images. Existing methods typically alleviate hallucinations by intervening in the attention score or output logits, focusing on "how to intervene" but overlooking the prerequisite "when to intervene", which leads to the "over-intervention" problem and subsequently introduces new hallucinations and unnecessary computational overhead. To address this gap, we first investigate the mechanism of visual neglect and reveal it can be accurately detected via head-level activation patterns in MLLMs. We thus propose V-ITI, a lightweight visual inference-time intervention framework integrating a Visual Neglect Detector that identifies visual neglect via head-level discriminative probes and a Visual Recall Intervenor that modulates activations with prestored visual activation information only when the visual neglect is detected. Extensive experiments across eight benchmarks and different MLLM families demonstrate that V-ITI consistently mitigates vision-related hallucinations while preserving general task performance.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03542
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
Sun, Nan
Zhang, Zhenyu
Lin, Xixun
Wang, Kun
Shang, Yanmin
Gu, Naibin
Wang, Shuohuan
Sun, Yu
Wu, Hua
Wang, Haifeng
Cao, Yanan
Computer Vision and Pattern Recognition
Artificial Intelligence
Multimodal Large Language Models (MLLMs) excel in numerous vision-language tasks yet suffer from hallucinations, producing content inconsistent with input visuals, that undermine reliability in precision-sensitive domains. This issue stems from a fundamental problem of visual neglect, where models fail to adequately prioritize input images. Existing methods typically alleviate hallucinations by intervening in the attention score or output logits, focusing on "how to intervene" but overlooking the prerequisite "when to intervene", which leads to the "over-intervention" problem and subsequently introduces new hallucinations and unnecessary computational overhead. To address this gap, we first investigate the mechanism of visual neglect and reveal it can be accurately detected via head-level activation patterns in MLLMs. We thus propose V-ITI, a lightweight visual inference-time intervention framework integrating a Visual Neglect Detector that identifies visual neglect via head-level discriminative probes and a Visual Recall Intervenor that modulates activations with prestored visual activation information only when the visual neglect is detected. Extensive experiments across eight benchmarks and different MLLM families demonstrate that V-ITI consistently mitigates vision-related hallucinations while preserving general task performance.
title V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.03542