MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Wenqian, Liu, Bohan, Zheng, Guangtao, Wang, Di, Ma, Yunsheng, Cao, Xu, Lai, Bolin, Rehg, James M., Zhang, Aidong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
NeuronTune: Towards Self-Guided Spurious Bias Mitigation
by: Zheng, Guangtao, et al.
Published: (2025)
by: Zheng, Guangtao, et al.
Published: (2025)
Benchmarking Spurious Bias in Few-Shot Image Classifiers
by: Zheng, Guangtao, et al.
Published: (2024)
by: Zheng, Guangtao, et al.
Published: (2024)
Spuriousness-Aware Meta-Learning for Learning Robust Classifiers
by: Zheng, Guangtao, et al.
Published: (2024)
by: Zheng, Guangtao, et al.
Published: (2024)
Improving Group Robustness on Spurious Correlation via Evidential Alignment
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
Learning Robust Classifiers with Self-Guided Spurious Correlation Mitigation
by: Zheng, Guangtao, et al.
Published: (2024)
by: Zheng, Guangtao, et al.
Published: (2024)
What is the Visual Cognition Gap between Humans and Multimodal LLMs?
by: Cao, Xu, et al.
Published: (2024)
by: Cao, Xu, et al.
Published: (2024)
ShortcutProbe: Probing Prediction Shortcuts for Learning Robust Models
by: Zheng, Guangtao, et al.
Published: (2025)
by: Zheng, Guangtao, et al.
Published: (2025)
Rectifying Shortcut Behaviors in Preference-based Reward Learning
by: Ye, Wenqian, et al.
Published: (2025)
by: Ye, Wenqian, et al.
Published: (2025)
The Clever Hans Mirage: A Comprehensive Survey on Spurious Correlations in Machine Learning
by: Ye, Wenqian, et al.
Published: (2024)
by: Ye, Wenqian, et al.
Published: (2024)
Towards Online Multi-Modal Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2025)
by: Li, Xinpeng, et al.
Published: (2025)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
by: Qian, Yusu, et al.
Published: (2024)
by: Qian, Yusu, et al.
Published: (2024)
Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
by: Lee, Sangmin, et al.
Published: (2024)
by: Lee, Sangmin, et al.
Published: (2024)
SocialGesture: Delving into Multi-person Gesture Understanding
by: Cao, Xu, et al.
Published: (2025)
by: Cao, Xu, et al.
Published: (2025)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2026)
by: Li, Xinpeng, et al.
Published: (2026)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
by: Lai, Bolin, et al.
Published: (2022)
by: Lai, Bolin, et al.
Published: (2022)
Towards Social AI: A Survey on Understanding Social Interactions
by: Lee, Sangmin, et al.
Published: (2024)
by: Lee, Sangmin, et al.
Published: (2024)
AdvST: Revisiting Data Augmentations for Single Domain Generalization
by: Zheng, Guangtao, et al.
Published: (2023)
by: Zheng, Guangtao, et al.
Published: (2023)
Learning Predictive Visuomotor Coordination
by: Jia, Wenqi, et al.
Published: (2025)
by: Jia, Wenqi, et al.
Published: (2025)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
by: Lai, Bolin, et al.
Published: (2023)
by: Lai, Bolin, et al.
Published: (2023)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA
by: Ye, Hanrong, et al.
Published: (2024)
by: Ye, Hanrong, et al.
Published: (2024)
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
by: Feng, Di, et al.
Published: (2025)
by: Feng, Di, et al.
Published: (2025)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
by: Daxberger, Erik, et al.
Published: (2025)
by: Daxberger, Erik, et al.
Published: (2025)
Towards Formalizing Spuriousness of Biased Datasets Using Partial Information Decomposition
by: Halder, Barproda, et al.
Published: (2024)
by: Halder, Barproda, et al.
Published: (2024)
LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs
by: Ma, Yunsheng, et al.
Published: (2023)
by: Ma, Yunsheng, et al.
Published: (2023)
InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding
by: Liu, Haogeng, et al.
Published: (2024)
by: Liu, Haogeng, et al.
Published: (2024)
MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
by: Ma, Haoxuan, et al.
Published: (2026)
by: Ma, Haoxuan, et al.
Published: (2026)
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
by: Lai, Bolin, et al.
Published: (2023)
by: Lai, Bolin, et al.
Published: (2023)
MM-IFEngine: Towards Multimodal Instruction Following
by: Ding, Shengyuan, et al.
Published: (2025)
by: Ding, Shengyuan, et al.
Published: (2025)
Towards Better Statistical Understanding of Watermarking LLMs
by: Cai, Zhongze, et al.
Published: (2024)
by: Cai, Zhongze, et al.
Published: (2024)
The Inscription in the ’Du khang of Dgung ’phur Monastery, Spu rang (Mnga’ ris)
by: Tropper, Kurt
Published: (2016)
by: Tropper, Kurt
Published: (2016)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
by: Lin, Sheng-Chieh, et al.
Published: (2024)
by: Lin, Sheng-Chieh, et al.
Published: (2024)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
by: Xiong, Guangzhi, et al.
Published: (2025)
by: Xiong, Guangzhi, et al.
Published: (2025)
The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory
by: Tang, Luoxi, et al.
Published: (2026)
by: Tang, Luoxi, et al.
Published: (2026)
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
by: Sinha, Sanchit, et al.
Published: (2026)
by: Sinha, Sanchit, et al.
Published: (2026)
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
by: Hosseini, Parsa, et al.
Published: (2025)
by: Hosseini, Parsa, et al.
Published: (2025)
Mixup Helps Understanding Multimodal Video Better
by: Ma, Xiaoyu, et al.
Published: (2025)
by: Ma, Xiaoyu, et al.
Published: (2025)
OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation
by: Jia, Xiaojun, et al.
Published: (2025)
by: Jia, Xiaojun, et al.
Published: (2025)
Similar Items
-
SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias
by: Ye, Wenqian, et al.
Published: (2025) -
NeuronTune: Towards Self-Guided Spurious Bias Mitigation
by: Zheng, Guangtao, et al.
Published: (2025) -
Benchmarking Spurious Bias in Few-Shot Image Classifiers
by: Zheng, Guangtao, et al.
Published: (2024) -
Spuriousness-Aware Meta-Learning for Learning Robust Classifiers
by: Zheng, Guangtao, et al.
Published: (2024) -
Improving Group Robustness on Spurious Correlation via Evidential Alignment
by: Ye, Wenqian, et al.
Published: (2025)