Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oh, Changdae, Fang, Zhen, Im, Shawn, Du, Xuefeng, Li, Yixuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908377490128896
author Oh, Changdae
Fang, Zhen
Im, Shawn
Du, Xuefeng
Li, Yixuan
author_facet Oh, Changdae
Fang, Zhen
Im, Shawn
Du, Xuefeng
Li, Yixuan
contents Multimodal large language models (MLLMs) have shown promising capabilities but struggle under distribution shifts, where evaluation data differ from instruction tuning distributions. Although previous works have provided empirical evaluations, we argue that establishing a formal framework that can characterize and quantify the risk of MLLMs is necessary to ensure the safe and reliable application of MLLMs in the real world. By taking an information-theoretic perspective, we propose the first theoretical framework that enables the quantification of the maximum risk of MLLMs under distribution shifts. Central to our framework is the introduction of Effective Mutual Information (EMI), a principled metric that quantifies the relevance between input queries and model responses. We derive an upper bound for the EMI difference between in-distribution (ID) and out-of-distribution (OOD) data, connecting it to visual and textual distributional discrepancies. Extensive experiments on real benchmark datasets, spanning 61 shift scenarios, empirically validate our theoretical insights.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00577
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
Oh, Changdae
Fang, Zhen
Im, Shawn
Du, Xuefeng
Li, Yixuan
Artificial Intelligence
Computation and Language
Machine Learning
Multimodal large language models (MLLMs) have shown promising capabilities but struggle under distribution shifts, where evaluation data differ from instruction tuning distributions. Although previous works have provided empirical evaluations, we argue that establishing a formal framework that can characterize and quantify the risk of MLLMs is necessary to ensure the safe and reliable application of MLLMs in the real world. By taking an information-theoretic perspective, we propose the first theoretical framework that enables the quantification of the maximum risk of MLLMs under distribution shifts. Central to our framework is the introduction of Effective Mutual Information (EMI), a principled metric that quantifies the relevance between input queries and model responses. We derive an upper bound for the EMI difference between in-distribution (ID) and out-of-distribution (OOD) data, connecting it to visual and textual distributional discrepancies. Extensive experiments on real benchmark datasets, spanning 61 shift scenarios, empirically validate our theoretical insights.
title Understanding Multimodal LLMs Under Distribution Shifts: An Information-Theoretic Approach
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.00577