Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Zhongxing, Wang, Zhonghua, Qian, Zhe, Shi, Dachuan, Tang, Feilong, Hu, Ming, Su, Shiyan, Zou, Xiaocheng, Feng, Wei, Mahapatra, Dwarikanath, Peng, Yifan, Lin, Mingquan, Ge, Zongyuan
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912965766152192
author Xu, Zhongxing
Wang, Zhonghua
Qian, Zhe
Shi, Dachuan
Tang, Feilong
Hu, Ming
Su, Shiyan
Zou, Xiaocheng
Feng, Wei
Mahapatra, Dwarikanath
Peng, Yifan
Lin, Mingquan
Ge, Zongyuan
author_facet Xu, Zhongxing
Wang, Zhonghua
Qian, Zhe
Shi, Dachuan
Tang, Feilong
Hu, Ming
Su, Shiyan
Zou, Xiaocheng
Feng, Wei
Mahapatra, Dwarikanath
Peng, Yifan
Lin, Mingquan
Ge, Zongyuan
contents Recent advancements in multimodal large reasoning models (MLRMs) have significantly improved performance in visual question answering. However, we observe that transition words (e.g., because, however, and wait) are closely associated with hallucinations and tend to exhibit high-entropy states. We argue that adequate contextual reasoning information can be directly extracted from the token probability distribution. Inspired by superposed representation theory, we propose leveraging latent superposed reasoning to integrate multiple candidate semantics and maintain latent reasoning trajectories. The hypothesis is that reliance on discrete textual inputs may drive the model toward sequential explicit reasoning, underutilizing dense contextual cues during high-entropy reasoning stages. Therefore, we propose constructing rich semantic representations from the token probability distributions to enhance in-context reasoning. With this goal, we present Latent Entropy-Aware Decoding (LEAD), an efficient plug-and-play decoding strategy that leverages semantic context to achieve reliable reasoning. The heart of our method lies in entropy-aware reasoning mode switching. The model employs probability-weighted continuous embeddings under high-entropy states and transitions back to discrete token embeddings as entropy decreases. Moreover, we propose a prior-guided visual anchor injection strategy that encourages the model to focus on visual information. Extensive experiments show that LEAD effectively mitigates hallucinations across various MLRMs on multiple benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_13366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
Xu, Zhongxing
Wang, Zhonghua
Qian, Zhe
Shi, Dachuan
Tang, Feilong
Hu, Ming
Su, Shiyan
Zou, Xiaocheng
Feng, Wei
Mahapatra, Dwarikanath
Peng, Yifan
Lin, Mingquan
Ge, Zongyuan
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advancements in multimodal large reasoning models (MLRMs) have significantly improved performance in visual question answering. However, we observe that transition words (e.g., because, however, and wait) are closely associated with hallucinations and tend to exhibit high-entropy states. We argue that adequate contextual reasoning information can be directly extracted from the token probability distribution. Inspired by superposed representation theory, we propose leveraging latent superposed reasoning to integrate multiple candidate semantics and maintain latent reasoning trajectories. The hypothesis is that reliance on discrete textual inputs may drive the model toward sequential explicit reasoning, underutilizing dense contextual cues during high-entropy reasoning stages. Therefore, we propose constructing rich semantic representations from the token probability distributions to enhance in-context reasoning. With this goal, we present Latent Entropy-Aware Decoding (LEAD), an efficient plug-and-play decoding strategy that leverages semantic context to achieve reliable reasoning. The heart of our method lies in entropy-aware reasoning mode switching. The model employs probability-weighted continuous embeddings under high-entropy states and transitions back to discrete token embeddings as entropy decreases. Moreover, we propose a prior-guided visual anchor injection strategy that encourages the model to focus on visual information. Extensive experiments show that LEAD effectively mitigates hallucinations across various MLRMs on multiple benchmarks.
title Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.13366