Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Ruoxi, Ma, Haoxuan, Hai, Zhengfei, Huang, Yiyan, Duan, Ranjie, Zhang, Tianle, Yang, Xu, Ye, Ziyi, Ma, Xingjun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918520919425024
author Cheng, Ruoxi
Ma, Haoxuan
Hai, Zhengfei
Huang, Yiyan
Duan, Ranjie
Zhang, Tianle
Yang, Xu
Ye, Ziyi
Ma, Xingjun
author_facet Cheng, Ruoxi
Ma, Haoxuan
Hai, Zhengfei
Huang, Yiyan
Duan, Ranjie
Zhang, Tianle
Yang, Xu
Ye, Ziyi
Ma, Xingjun
contents Large Vision-Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated content conflicts with visual facts. Existing mitigation methods either rely on costly external interventions, such as instruction tuning and retrieval, or use internal mechanisms that remain limited by flawed attention weights and entangled hidden representations. We propose Adversarial Orthogonal Disentanglement (AOD), a latent geometric framework for mitigating LVLM hallucinations. AOD learns a hallucination-related direction through a minimax objective: a classifier concentrates hallucination signals into the projected component, while an adversary removes them from the orthogonal residual space via a Gradient Reversal Layer. The learned direction enables a training-free dual-forward-pass contrastive decoding strategy that suppresses hallucinations while preserving general capabilities. Experiments on three LVLMs across four hallucination and four utility benchmarks show that AOD consistently outperforms strong baselines. It improves POPE accuracy by over 6\% on average, boosts AMBER by 6\%, and maintains strong performance on utility tasks such as MMMU. Further analysis shows robust transfer across datasets, suggesting that AOD captures general hallucination-related biases rather than dataset-specific artifacts. Our source code and datasets are available at https://github.com/Hunter-Wrynn/AOD.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25377
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
Cheng, Ruoxi
Ma, Haoxuan
Hai, Zhengfei
Huang, Yiyan
Duan, Ranjie
Zhang, Tianle
Yang, Xu
Ye, Ziyi
Ma, Xingjun
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision-Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated content conflicts with visual facts. Existing mitigation methods either rely on costly external interventions, such as instruction tuning and retrieval, or use internal mechanisms that remain limited by flawed attention weights and entangled hidden representations. We propose Adversarial Orthogonal Disentanglement (AOD), a latent geometric framework for mitigating LVLM hallucinations. AOD learns a hallucination-related direction through a minimax objective: a classifier concentrates hallucination signals into the projected component, while an adversary removes them from the orthogonal residual space via a Gradient Reversal Layer. The learned direction enables a training-free dual-forward-pass contrastive decoding strategy that suppresses hallucinations while preserving general capabilities. Experiments on three LVLMs across four hallucination and four utility benchmarks show that AOD consistently outperforms strong baselines. It improves POPE accuracy by over 6\% on average, boosts AMBER by 6\%, and maintains strong performance on utility tasks such as MMMU. Further analysis shows robust transfer across datasets, suggesting that AOD captures general hallucination-related biases rather than dataset-specific artifacts. Our source code and datasets are available at https://github.com/Hunter-Wrynn/AOD.
title Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.25377