Hallucination-aware intermediate representation edit in large vision-language models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Suo, Wei, Zhang, Hanzu, Zhang, Lijun, Ma, Ji, Wang, Peng, Zhang, Yanning
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915901212721152
author Suo, Wei
Zhang, Hanzu
Zhang, Lijun
Ma, Ji
Wang, Peng
Zhang, Yanning
author_facet Suo, Wei
Zhang, Hanzu
Zhang, Lijun
Ma, Ji
Wang, Peng
Zhang, Yanning
contents Large Vision-Language Models have demonstrated exceptional performance in multimodal reasoning and complex scene understanding. However, these models still face significant hallucination issues, where outputs contradict visual facts. Recent research on hallucination mitigation has focused on retraining methods and Contrastive Decoding (CD) methods. While both methods perform well, retraining methods require substantial training resources, and CD methods introduce dual inference overhead. These factors hinder their practical applicability. To address the above issue, we propose a framework for dynamically detecting hallucination representations and performing hallucination-eliminating edits on these representations. With minimal additional computational cost, we achieve state-of-the-art performance on existing benchmarks. Extensive experiments demonstrate the effectiveness of our approach, highlighting its efficient and robust hallucination elimination capability and its powerful controllability over hallucinations. Code is available at https://github.com/ASGO-MM/HIRE
format Preprint
id arxiv_https___arxiv_org_abs_2603_29405
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hallucination-aware intermediate representation edit in large vision-language models
Suo, Wei
Zhang, Hanzu
Zhang, Lijun
Ma, Ji
Wang, Peng
Zhang, Yanning
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision-Language Models have demonstrated exceptional performance in multimodal reasoning and complex scene understanding. However, these models still face significant hallucination issues, where outputs contradict visual facts. Recent research on hallucination mitigation has focused on retraining methods and Contrastive Decoding (CD) methods. While both methods perform well, retraining methods require substantial training resources, and CD methods introduce dual inference overhead. These factors hinder their practical applicability. To address the above issue, we propose a framework for dynamically detecting hallucination representations and performing hallucination-eliminating edits on these representations. With minimal additional computational cost, we achieve state-of-the-art performance on existing benchmarks. Extensive experiments demonstrate the effectiveness of our approach, highlighting its efficient and robust hallucination elimination capability and its powerful controllability over hallucinations. Code is available at https://github.com/ASGO-MM/HIRE
title Hallucination-aware intermediate representation edit in large vision-language models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.29405