FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhattacharya, Debarpan, Kulkarni, Apoorva, Ganapathy, Sriram
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917233920311296
author Bhattacharya, Debarpan
Kulkarni, Apoorva
Ganapathy, Sriram
author_facet Bhattacharya, Debarpan
Kulkarni, Apoorva
Ganapathy, Sriram
contents The accurate trust assessment of multimodal large language models (MLLMs) generated predictions, which can enable selective prediction and improve user confidence, is challenging due to the diverse multi-modal input paradigms. We propose Functionally Equivalent Sampling for Trust Assessment (FESTA), a multimodal input sampling technique for MLLMs, that generates an uncertainty measure based on the equivalent and complementary input samplings. The proposed task-preserving sampling approach for uncertainty quantification expands the input space to probe the consistency (through equivalent samples) and sensitivity (through complementary samples) of the model. FESTA uses only input-output access of the model (black-box), and does not require ground truth (unsupervised). The experiments are conducted with various off-the-shelf multi-modal LLMs, on both visual and audio reasoning tasks. The proposed FESTA uncertainty estimate achieves significant improvement (33.3% relative improvement for vision-LLMs and 29.6% relative improvement for audio-LLMs) in selective prediction performance, based on area-under-receiver-operating-characteristic curve (AUROC) metric in detecting mispredictions. The code implementation is open-sourced.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16648
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
Bhattacharya, Debarpan
Kulkarni, Apoorva
Ganapathy, Sriram
Artificial Intelligence
Computation and Language
Machine Learning
The accurate trust assessment of multimodal large language models (MLLMs) generated predictions, which can enable selective prediction and improve user confidence, is challenging due to the diverse multi-modal input paradigms. We propose Functionally Equivalent Sampling for Trust Assessment (FESTA), a multimodal input sampling technique for MLLMs, that generates an uncertainty measure based on the equivalent and complementary input samplings. The proposed task-preserving sampling approach for uncertainty quantification expands the input space to probe the consistency (through equivalent samples) and sensitivity (through complementary samples) of the model. FESTA uses only input-output access of the model (black-box), and does not require ground truth (unsupervised). The experiments are conducted with various off-the-shelf multi-modal LLMs, on both visual and audio reasoning tasks. The proposed FESTA uncertainty estimate achieves significant improvement (33.3% relative improvement for vision-LLMs and 29.6% relative improvement for audio-LLMs) in selective prediction performance, based on area-under-receiver-operating-characteristic curve (AUROC) metric in detecting mispredictions. The code implementation is open-sourced.
title FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.16648