VCE: A zero-cost hallucination mitigation method of LVLMs via visual contrastive editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yanbin, Li, Yisen, Tie, Guiyao, Qu, Xiaoye, Zhou, Pan, Wang, Hongfei, Zou, Zhaofan, Sun, Hao, Li, Xuelong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910153619537920
author Huang, Yanbin
Li, Yisen
Tie, Guiyao
Qu, Xiaoye
Zhou, Pan
Wang, Hongfei
Zou, Zhaofan
Sun, Hao
Li, Xuelong
author_facet Huang, Yanbin
Li, Yisen
Tie, Guiyao
Qu, Xiaoye
Zhou, Pan
Wang, Hongfei
Zou, Zhaofan
Sun, Hao
Li, Xuelong
contents Large vision-language models (LVLMs) frequently suffer from Object Hallucination (OH), wherein they generate descriptions containing objects that are not actually present in the input image. This phenomenon is particularly problematic in real-world applications such as medical imaging and autonomous driving, where accuracy is critical. Recent studies suggest that the hallucination problem may stem from language priors: biases learned during pretraining that cause LVLMs to generate words based on their statistical co-occurrence. To mitigate this problem, we propose Visual Contrastive Editing (VCE), a novel post-hoc method that identifies and suppresses hallucinatory tendencies by analyzing the model's response to contrastive visual perturbations. Using Singular Value Decomposition (SVD), we decompose the model's activation patterns to isolate hallucination subspaces and apply targeted parameter edits to attenuate its influence. Unlike existing approaches that require fine-tuning or labeled data, VCE operates as a label-free intervention, making it both scalable and practical for deployment in resource-constrained settings. Experimental results demonstrate that VCE effectively reduces object hallucination across multiple benchmarks while maintaining the model's original computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2604_19412
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VCE: A zero-cost hallucination mitigation method of LVLMs via visual contrastive editing
Huang, Yanbin
Li, Yisen
Tie, Guiyao
Qu, Xiaoye
Zhou, Pan
Wang, Hongfei
Zou, Zhaofan
Sun, Hao
Li, Xuelong
Computer Vision and Pattern Recognition
Computation and Language
Large vision-language models (LVLMs) frequently suffer from Object Hallucination (OH), wherein they generate descriptions containing objects that are not actually present in the input image. This phenomenon is particularly problematic in real-world applications such as medical imaging and autonomous driving, where accuracy is critical. Recent studies suggest that the hallucination problem may stem from language priors: biases learned during pretraining that cause LVLMs to generate words based on their statistical co-occurrence. To mitigate this problem, we propose Visual Contrastive Editing (VCE), a novel post-hoc method that identifies and suppresses hallucinatory tendencies by analyzing the model's response to contrastive visual perturbations. Using Singular Value Decomposition (SVD), we decompose the model's activation patterns to isolate hallucination subspaces and apply targeted parameter edits to attenuate its influence. Unlike existing approaches that require fine-tuning or labeled data, VCE operates as a label-free intervention, making it both scalable and practical for deployment in resource-constrained settings. Experimental results demonstrate that VCE effectively reduces object hallucination across multiple benchmarks while maintaining the model's original computational efficiency.
title VCE: A zero-cost hallucination mitigation method of LVLMs via visual contrastive editing
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2604.19412