Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fang, Yixiong, Yang, Ziran, Chen, Zhaorun, Zhao, Zhuokai, Zhou, Jiawei
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917170765627392
author Fang, Yixiong
Yang, Ziran
Chen, Zhaorun
Zhao, Zhuokai
Zhou, Jiawei
author_facet Fang, Yixiong
Yang, Ziran
Chen, Zhaorun
Zhao, Zhuokai
Zhou, Jiawei
contents Large vision-language models (LVLMs) excel at multimodal tasks but are prone to misinterpreting visual inputs, often resulting in hallucinations and unreliable outputs. We present DROPOUT DECODING, a novel inference-time approach that quantifies the uncertainty of visual tokens and selectively masks uncertain tokens to improve decoding. Our method measures the uncertainty of each visual token by projecting it onto the text space and decomposing it into aleatoric and epistemic components. Specifically, we focus on epistemic uncertainty, which captures perception-related errors more effectively. Inspired by dropout regularization, we introduce uncertainty-guided token dropout, which applies the dropout principle to input visual tokens instead of model parameters, and during inference rather than training. By aggregating predictions from an ensemble of masked decoding contexts, we can robustly mitigate errors arising from visual token misinterpretations. Evaluations on benchmarks including CHAIR, THRONE, and MMBench demonstrate that DROPOUT DECODING significantly reduces object hallucinations (OH) and enhances both reliability and quality of LVLM outputs across diverse visual contexts. Code is released at https://github.com/kigb/DropoutDecoding.
format Preprint
id arxiv_https___arxiv_org_abs_2412_06474
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
Fang, Yixiong
Yang, Ziran
Chen, Zhaorun
Zhao, Zhuokai
Zhou, Jiawei
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Large vision-language models (LVLMs) excel at multimodal tasks but are prone to misinterpreting visual inputs, often resulting in hallucinations and unreliable outputs. We present DROPOUT DECODING, a novel inference-time approach that quantifies the uncertainty of visual tokens and selectively masks uncertain tokens to improve decoding. Our method measures the uncertainty of each visual token by projecting it onto the text space and decomposing it into aleatoric and epistemic components. Specifically, we focus on epistemic uncertainty, which captures perception-related errors more effectively. Inspired by dropout regularization, we introduce uncertainty-guided token dropout, which applies the dropout principle to input visual tokens instead of model parameters, and during inference rather than training. By aggregating predictions from an ensemble of masked decoding contexts, we can robustly mitigate errors arising from visual token misinterpretations. Evaluations on benchmarks including CHAIR, THRONE, and MMBench demonstrate that DROPOUT DECODING significantly reduces object hallucinations (OH) and enhances both reliability and quality of LVLM outputs across diverse visual contexts. Code is released at https://github.com/kigb/DropoutDecoding.
title Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.06474