Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Jinlin, Wang, Yuran, Yuan, Yifei, Zhou, Xiao, Zhang, Yingying, Yong, Xixian, Zheng, Yefeng, Wu, Xian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908604450209792
author Li, Jinlin
Wang, Yuran
Yuan, Yifei
Zhou, Xiao
Zhang, Yingying
Yong, Xixian
Zheng, Yefeng
Wu, Xian
author_facet Li, Jinlin
Wang, Yuran
Yuan, Yifei
Zhou, Xiao
Zhang, Yingying
Yong, Xixian
Zheng, Yefeng
Wu, Xian
contents Large Vision-Language Models (LVLMs) have recently achieved impressive results in multimodal tasks such as image captioning and visual question answering. However, they remain prone to object hallucination -- generating descriptions of nonexistent or misidentified objects. Prior work has partially mitigated this via auxiliary training objectives or external modules, but challenges remain in terms of scalability, adaptability, and model independence. To address these limitations, we propose Adaptive Token Ensemble Decoding (ATED), a training-free, token-level ensemble framework that mitigates hallucination by aggregating predictions from multiple LVLMs during inference. ATED dynamically computes uncertainty-based weights for each model, reflecting their reliability at each decoding step. It also integrates diverse decoding paths to improve contextual grounding and semantic consistency. Experiments on standard hallucination detection benchmarks demonstrate that ATED significantly outperforms state-of-the-art methods, reducing hallucination without compromising fluency or relevance. Our findings highlight the benefits of adaptive ensembling and point to a promising direction for improving LVLM robustness in high-stakes applications. The code is available at https://github.com/jinlin2021/ATED.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18321
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding
Li, Jinlin
Wang, Yuran
Yuan, Yifei
Zhou, Xiao
Zhang, Yingying
Yong, Xixian
Zheng, Yefeng
Wu, Xian
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) have recently achieved impressive results in multimodal tasks such as image captioning and visual question answering. However, they remain prone to object hallucination -- generating descriptions of nonexistent or misidentified objects. Prior work has partially mitigated this via auxiliary training objectives or external modules, but challenges remain in terms of scalability, adaptability, and model independence. To address these limitations, we propose Adaptive Token Ensemble Decoding (ATED), a training-free, token-level ensemble framework that mitigates hallucination by aggregating predictions from multiple LVLMs during inference. ATED dynamically computes uncertainty-based weights for each model, reflecting their reliability at each decoding step. It also integrates diverse decoding paths to improve contextual grounding and semantic consistency. Experiments on standard hallucination detection benchmarks demonstrate that ATED significantly outperforms state-of-the-art methods, reducing hallucination without compromising fluency or relevance. Our findings highlight the benefits of adaptive ensembling and point to a promising direction for improving LVLM robustness in high-stakes applications. The code is available at https://github.com/jinlin2021/ATED.
title Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.18321