Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Ji, Suo, Wei, Wang, Peng, Zhang, Yanning
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914429213343744
author Ma, Ji
Suo, Wei
Wang, Peng
Zhang, Yanning
author_facet Ma, Ji
Suo, Wei
Wang, Peng
Zhang, Yanning
contents Multimodal Chain-of-Thought (MCoT) models have demonstrated impressive capability in complex visual reasoning tasks. Unfortunately, recent studies reveal that they suffer from severe hallucination problems due to diminished visual attention during the generation process. However, visual attention decay is a well-studied problem in Large Vision-Language Models (LVLMs). Considering the fundamental differences in reasoning processes between MCoT models and traditional LVLMs, we raise a basic question: Whether MCoT models have unique causes of hallucinations? To answer this question, we systematically investigate the hallucination patterns of MCoT models and find that fabricated texts are primarily generated in associative reasoning steps, which we term divergent thinking. Leveraging these insights, we introduce a simple yet effective strategy that can effectively localize divergent thinking steps and intervene in the decoding process to mitigate hallucinations. Extensive experiments show that our method outperforms existing methods by a large margin. More importantly, our proposed method can be conveniently integrated with other hallucination mitigation methods and further boost their performance. The code is publicly available at https://github.com/ASGO-MM/MCoT-hallucination.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27201
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
Ma, Ji
Suo, Wei
Wang, Peng
Zhang, Yanning
Computer Vision and Pattern Recognition
Multimodal Chain-of-Thought (MCoT) models have demonstrated impressive capability in complex visual reasoning tasks. Unfortunately, recent studies reveal that they suffer from severe hallucination problems due to diminished visual attention during the generation process. However, visual attention decay is a well-studied problem in Large Vision-Language Models (LVLMs). Considering the fundamental differences in reasoning processes between MCoT models and traditional LVLMs, we raise a basic question: Whether MCoT models have unique causes of hallucinations? To answer this question, we systematically investigate the hallucination patterns of MCoT models and find that fabricated texts are primarily generated in associative reasoning steps, which we term divergent thinking. Leveraging these insights, we introduce a simple yet effective strategy that can effectively localize divergent thinking steps and intervene in the decoding process to mitigate hallucinations. Extensive experiments show that our method outperforms existing methods by a large margin. More importantly, our proposed method can be conveniently integrated with other hallucination mitigation methods and further boost their performance. The code is publicly available at https://github.com/ASGO-MM/MCoT-hallucination.
title Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.27201