Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: You, Liangliang, Yao, Junchi, Yang, Shu, Hu, Guimin, Hu, Lijie, Wang, Di
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912418965225472
author You, Liangliang
Yao, Junchi
Yang, Shu
Hu, Guimin
Hu, Lijie
Wang, Di
author_facet You, Liangliang
Yao, Junchi
Yang, Shu
Hu, Guimin
Hu, Lijie
Wang, Di
contents While multimodal large language models excel at various tasks, they still suffer from hallucinations, which limit their reliability and scalability for broader domain applications. To address this issue, recent research mainly focuses on objective hallucination. However, for sequential images, besides objective hallucination, there is also behavioral hallucination, which is less studied. This work aims to fill in the gap. We first reveal that behavioral hallucinations mainly arise from two key factors: prior-driven bias and the snowball effect. Based on these observations, we introduce SHE (Sequence Hallucination Eradication), a lightweight, two-stage framework that (1) detects hallucinations via visual-textual alignment check using our proposed adaptive temporal window and (2) mitigates them via orthogonal projection onto the joint embedding space. We also propose a new metric (BEACH) to quantify behavioral hallucination severity. Empirical results on standard benchmarks demonstrate that SHE reduces behavioral hallucination by over 10% on BEACH while maintaining descriptive accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07184
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
You, Liangliang
Yao, Junchi
Yang, Shu
Hu, Guimin
Hu, Lijie
Wang, Di
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
While multimodal large language models excel at various tasks, they still suffer from hallucinations, which limit their reliability and scalability for broader domain applications. To address this issue, recent research mainly focuses on objective hallucination. However, for sequential images, besides objective hallucination, there is also behavioral hallucination, which is less studied. This work aims to fill in the gap. We first reveal that behavioral hallucinations mainly arise from two key factors: prior-driven bias and the snowball effect. Based on these observations, we introduce SHE (Sequence Hallucination Eradication), a lightweight, two-stage framework that (1) detects hallucinations via visual-textual alignment check using our proposed adaptive temporal window and (2) mitigates them via orthogonal projection onto the joint embedding space. We also propose a new metric (BEACH) to quantify behavioral hallucination severity. Empirical results on standard benchmarks demonstrate that SHE reduces behavioral hallucination by over 10% on BEACH while maintaining descriptive accuracy.
title Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.07184