The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Azad, Asif, Hossain, Mohammad Sadat, Shanto, MD Sadik Hossain, Rahman, M Saifur, Parvez, Md Rizwan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918304262651904
author Azad, Asif
Hossain, Mohammad Sadat
Shanto, MD Sadik Hossain
Rahman, M Saifur
Parvez, Md Rizwan
author_facet Azad, Asif
Hossain, Mohammad Sadat
Shanto, MD Sadik Hossain
Rahman, M Saifur
Parvez, Md Rizwan
contents Vision-Language Models (VLMs) have achieved remarkable progress in complex visual understanding across scientific and reasoning tasks. While performance benchmarking has advanced our understanding of these capabilities, the critical dimension of uncertainty quantification has received insufficient attention. Therefore, unlike prior conformal prediction studies that focused on limited settings, we conduct a comprehensive uncertainty benchmarking study, evaluating 18 state-of-the-art VLMs (open and closed-source) across 6 multimodal datasets with 3 distinct scoring functions. For closed-source models lacking token-level logprob access, we develop and validate instruction-guided likelihood proxies. Our findings demonstrate that larger models consistently exhibit better uncertainty quantification; models that know more also know better what they don't know. More certain models achieve higher accuracy, while mathematical and reasoning tasks elicit poorer uncertainty performance across all models compared to other domains. This work establishes a foundation for reliable uncertainty evaluation in multimodal systems.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
Azad, Asif
Hossain, Mohammad Sadat
Shanto, MD Sadik Hossain
Rahman, M Saifur
Parvez, Md Rizwan
Artificial Intelligence
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) have achieved remarkable progress in complex visual understanding across scientific and reasoning tasks. While performance benchmarking has advanced our understanding of these capabilities, the critical dimension of uncertainty quantification has received insufficient attention. Therefore, unlike prior conformal prediction studies that focused on limited settings, we conduct a comprehensive uncertainty benchmarking study, evaluating 18 state-of-the-art VLMs (open and closed-source) across 6 multimodal datasets with 3 distinct scoring functions. For closed-source models lacking token-level logprob access, we develop and validate instruction-guided likelihood proxies. Our findings demonstrate that larger models consistently exhibit better uncertainty quantification; models that know more also know better what they don't know. More certain models achieve higher accuracy, while mathematical and reasoning tasks elicit poorer uncertainty performance across all models compared to other domains. This work establishes a foundation for reliable uncertainty evaluation in multimodal systems.
title The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
topic Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.13379