Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Padhi, Trilok, Kaur, Ramneet, Cobb, Adam D., Acharya, Manoj, Roy, Anirban, Samplawski, Colin, Matejek, Brian, Berenbeim, Alexander M., Bastian, Nathaniel D., Jha, Susmit
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916723330908160
author Padhi, Trilok
Kaur, Ramneet
Cobb, Adam D.
Acharya, Manoj
Roy, Anirban
Samplawski, Colin
Matejek, Brian
Berenbeim, Alexander M.
Bastian, Nathaniel D.
Jha, Susmit
author_facet Padhi, Trilok
Kaur, Ramneet
Cobb, Adam D.
Acharya, Manoj
Roy, Anirban
Samplawski, Colin
Matejek, Brian
Berenbeim, Alexander M.
Bastian, Nathaniel D.
Jha, Susmit
contents We introduce a novel approach for calibrating uncertainty quantification (UQ) tailored for multi-modal large language models (LLMs). Existing state-of-the-art UQ methods rely on consistency among multiple responses generated by the LLM on an input query under diverse settings. However, these approaches often report higher confidence in scenarios where the LLM is consistently incorrect. This leads to a poorly calibrated confidence with respect to accuracy. To address this, we leverage cross-modal consistency in addition to self-consistency to improve the calibration of the multi-modal models. Specifically, we ground the textual responses to the visual inputs. The confidence from the grounding model is used to calibrate the overall confidence. Given that using a grounding model adds its own uncertainty in the pipeline, we apply temperature scaling - a widely accepted parametric calibration technique - to calibrate the grounding model's confidence in the accuracy of generated responses. We evaluate the proposed approach across multiple multi-modal tasks, such as medical question answering (Slake) and visual question answering (VQAv2), considering multi-modal models such as LLaVA-Med and LLaVA. The experiments demonstrate that the proposed framework achieves significantly improved calibration on both tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03788
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
Padhi, Trilok
Kaur, Ramneet
Cobb, Adam D.
Acharya, Manoj
Roy, Anirban
Samplawski, Colin
Matejek, Brian
Berenbeim, Alexander M.
Bastian, Nathaniel D.
Jha, Susmit
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
We introduce a novel approach for calibrating uncertainty quantification (UQ) tailored for multi-modal large language models (LLMs). Existing state-of-the-art UQ methods rely on consistency among multiple responses generated by the LLM on an input query under diverse settings. However, these approaches often report higher confidence in scenarios where the LLM is consistently incorrect. This leads to a poorly calibrated confidence with respect to accuracy. To address this, we leverage cross-modal consistency in addition to self-consistency to improve the calibration of the multi-modal models. Specifically, we ground the textual responses to the visual inputs. The confidence from the grounding model is used to calibrate the overall confidence. Given that using a grounding model adds its own uncertainty in the pipeline, we apply temperature scaling - a widely accepted parametric calibration technique - to calibrate the grounding model's confidence in the accuracy of generated responses. We evaluate the proposed approach across multiple multi-modal tasks, such as medical question answering (Slake) and visual question answering (VQAv2), considering multi-modal models such as LLaVA-Med and LLaVA. The experiments demonstrate that the proposed framework achieves significantly improved calibration on both tasks.
title Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.03788