Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Zijun, Hu, Wenbo, He, Guande, Deng, Zhijie, Zhang, Zheng, Hong, Richang
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910764139282432
author Chen, Zijun
Hu, Wenbo
He, Guande
Deng, Zhijie
Zhang, Zheng
Hong, Richang
author_facet Chen, Zijun
Hu, Wenbo
He, Guande
Deng, Zhijie
Zhang, Zheng
Hong, Richang
contents Multimodal large language models (MLLMs) combine visual and textual data for tasks such as image captioning and visual question answering. Proper uncertainty calibration is crucial, yet challenging, for reliable use in areas like healthcare and autonomous driving. This paper investigates representative MLLMs, focusing on their calibration across various scenarios, including before and after visual fine-tuning, as well as before and after multimodal training of the base LLMs. We observed miscalibration in their performance, and at the same time, no significant differences in calibration across these scenarios. We also highlight how uncertainty differs between text and images and how their integration affects overall uncertainty. To better understand MLLMs' miscalibration and their ability to self-assess uncertainty, we construct the IDK (I don't know) dataset, which is key to evaluating how they handle unknowns. Our findings reveal that MLLMs tend to give answers rather than admit uncertainty, but this self-assessment improves with proper prompt adjustments. Finally, to calibrate MLLMs and enhance model reliability, we propose techniques such as temperature scaling and iterative prompt optimization. Our results provide insights into improving MLLMs for effective and responsible deployment in multimodal applications. Code and IDK dataset: https://github.com/hfutml/Calibration-MLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14660
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models
Chen, Zijun
Hu, Wenbo
He, Guande
Deng, Zhijie
Zhang, Zheng
Hong, Richang
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimodal large language models (MLLMs) combine visual and textual data for tasks such as image captioning and visual question answering. Proper uncertainty calibration is crucial, yet challenging, for reliable use in areas like healthcare and autonomous driving. This paper investigates representative MLLMs, focusing on their calibration across various scenarios, including before and after visual fine-tuning, as well as before and after multimodal training of the base LLMs. We observed miscalibration in their performance, and at the same time, no significant differences in calibration across these scenarios. We also highlight how uncertainty differs between text and images and how their integration affects overall uncertainty. To better understand MLLMs' miscalibration and their ability to self-assess uncertainty, we construct the IDK (I don't know) dataset, which is key to evaluating how they handle unknowns. Our findings reveal that MLLMs tend to give answers rather than admit uncertainty, but this self-assessment improves with proper prompt adjustments. Finally, to calibrate MLLMs and enhance model reliability, we propose techniques such as temperature scaling and iterative prompt optimization. Our results provide insights into improving MLLMs for effective and responsible deployment in multimodal applications. Code and IDK dataset: https://github.com/hfutml/Calibration-MLLM.
title Unveiling Uncertainty: A Deep Dive into Calibration and Performance of Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.14660