SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hosseini, Parsa, Nawathe, Sumit, Moayeri, Mazda, Balasubramanian, Sriram, Feizi, Soheil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915673600425984
author Hosseini, Parsa
Nawathe, Sumit
Moayeri, Mazda
Balasubramanian, Sriram
Feizi, Soheil
author_facet Hosseini, Parsa
Nawathe, Sumit
Moayeri, Mazda
Balasubramanian, Sriram
Feizi, Soheil
contents Unimodal vision models are known to rely on spurious correlations, but it remains unclear to what extent Multimodal Large Language Models (MLLMs) exhibit similar biases despite language supervision. In this paper, we investigate spurious bias in MLLMs and introduce SpurLens, a pipeline that leverages GPT-4 and open-set object detectors to automatically identify spurious visual cues without human supervision. Our findings reveal that spurious correlations cause two major failure modes in MLLMs: (1) over-reliance on spurious cues for object recognition, where removing these cues reduces accuracy, and (2) object hallucination, where spurious cues amplify the hallucination by over 10x. We validate our findings in various MLLMs and datasets. Beyond diagnosing these failures, we explore potential mitigation strategies, such as prompt ensembling and reasoning-based prompting, and conduct ablation studies to examine the root causes of spurious bias in MLLMs. By exposing the persistence of spurious correlations, our study calls for more rigorous evaluation methods and mitigation strategies to enhance the reliability of MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_08884
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
Hosseini, Parsa
Nawathe, Sumit
Moayeri, Mazda
Balasubramanian, Sriram
Feizi, Soheil
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Unimodal vision models are known to rely on spurious correlations, but it remains unclear to what extent Multimodal Large Language Models (MLLMs) exhibit similar biases despite language supervision. In this paper, we investigate spurious bias in MLLMs and introduce SpurLens, a pipeline that leverages GPT-4 and open-set object detectors to automatically identify spurious visual cues without human supervision. Our findings reveal that spurious correlations cause two major failure modes in MLLMs: (1) over-reliance on spurious cues for object recognition, where removing these cues reduces accuracy, and (2) object hallucination, where spurious cues amplify the hallucination by over 10x. We validate our findings in various MLLMs and datasets. Beyond diagnosing these failures, we explore potential mitigation strategies, such as prompt ensembling and reasoning-based prompting, and conduct ablation studies to examine the root causes of spurious bias in MLLMs. By exposing the persistence of spurious correlations, our study calls for more rigorous evaluation methods and mitigation strategies to enhance the reliability of MLLMs.
title SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2503.08884