Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Meiqi, Cao, Yixin, Zhang, Yan, Lu, Chaochao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CELLO: Causal Evaluation of Large Vision-Language Models
di: Chen, Meiqi, et al.
Pubblicazione: (2024)
di: Chen, Meiqi, et al.
Pubblicazione: (2024)
Detecting Offensive Memes with Social Biases in Singapore Context Using Multimodal Large Language Models
di: Yuxuan, Cao, et al.
Pubblicazione: (2025)
di: Yuxuan, Cao, et al.
Pubblicazione: (2025)
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
di: Li, Zhuowan, et al.
Pubblicazione: (2022)
di: Li, Zhuowan, et al.
Pubblicazione: (2022)
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
di: Wang, Weihang, et al.
Pubblicazione: (2025)
di: Wang, Weihang, et al.
Pubblicazione: (2025)
Studying and Mitigating Biases in Sign Language Understanding Models
di: Atwell, Katherine, et al.
Pubblicazione: (2024)
di: Atwell, Katherine, et al.
Pubblicazione: (2024)
ADAM: An Embodied Causal Agent in Open-World Environments
di: Yu, Shu, et al.
Pubblicazione: (2024)
di: Yu, Shu, et al.
Pubblicazione: (2024)
Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
di: Weng, Fenghua, et al.
Pubblicazione: (2025)
Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation
di: Wang, Zehao, et al.
Pubblicazione: (2024)
di: Wang, Zehao, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
di: Wu, Jiulong, et al.
Pubblicazione: (2025)
di: Wu, Jiulong, et al.
Pubblicazione: (2025)
EFUF: Efficient Fine-grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models
di: Xing, Shangyu, et al.
Pubblicazione: (2024)
di: Xing, Shangyu, et al.
Pubblicazione: (2024)
Mitigating Object Hallucination via Robust Local Perception Search
di: Gao, Zixian, et al.
Pubblicazione: (2025)
di: Gao, Zixian, et al.
Pubblicazione: (2025)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
di: Zheng, Kening, et al.
Pubblicazione: (2024)
di: Zheng, Kening, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
di: Lu, Yifan, et al.
Pubblicazione: (2025)
di: Lu, Yifan, et al.
Pubblicazione: (2025)
CauSight: Learning to Supersense for Visual Causal Discovery
di: Zhang, Yize, et al.
Pubblicazione: (2025)
di: Zhang, Yize, et al.
Pubblicazione: (2025)
Learning Domain Knowledge in Multimodal Large Language Models through Reinforcement Fine-Tuning
di: Cao, Qinglong, et al.
Pubblicazione: (2026)
di: Cao, Qinglong, et al.
Pubblicazione: (2026)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
di: Chen, Boqi, et al.
Pubblicazione: (2026)
di: Chen, Boqi, et al.
Pubblicazione: (2026)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
di: Chang, Yue, et al.
Pubblicazione: (2024)
di: Chang, Yue, et al.
Pubblicazione: (2024)
Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
di: Yue, Zihao, et al.
Pubblicazione: (2024)
di: Yue, Zihao, et al.
Pubblicazione: (2024)
Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
di: Xu, Shilin, et al.
Pubblicazione: (2025)
di: Xu, Shilin, et al.
Pubblicazione: (2025)
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
di: Wang, Weiyun, et al.
Pubblicazione: (2024)
di: Wang, Weiyun, et al.
Pubblicazione: (2024)
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
di: Ma, Ruiqi, et al.
Pubblicazione: (2025)
di: Ma, Ruiqi, et al.
Pubblicazione: (2025)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
di: Zhong, Weihong, et al.
Pubblicazione: (2024)
Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
di: You, Liangliang, et al.
Pubblicazione: (2025)
di: You, Liangliang, et al.
Pubblicazione: (2025)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
di: Xu, Shilin, et al.
Pubblicazione: (2024)
di: Xu, Shilin, et al.
Pubblicazione: (2024)
Mitigating Object Hallucination via Concentric Causal Attention
di: Xing, Yun, et al.
Pubblicazione: (2024)
di: Xing, Yun, et al.
Pubblicazione: (2024)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
di: Xia, Haotian, et al.
Pubblicazione: (2024)
di: Xia, Haotian, et al.
Pubblicazione: (2024)
Model Composition for Multimodal Large Language Models
di: Chen, Chi, et al.
Pubblicazione: (2024)
di: Chen, Chi, et al.
Pubblicazione: (2024)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
di: Wu, Yuhang, et al.
Pubblicazione: (2024)
di: Wu, Yuhang, et al.
Pubblicazione: (2024)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
di: Yuan, Qianhao, et al.
Pubblicazione: (2025)
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
di: Du, Yiyang, et al.
Pubblicazione: (2025)
di: Du, Yiyang, et al.
Pubblicazione: (2025)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
di: Zhu, Yingjie, et al.
Pubblicazione: (2025)
di: Zhu, Yingjie, et al.
Pubblicazione: (2025)
Strengthening Multimodal Large Language Model with Bootstrapped Preference Optimization
di: Pi, Renjie, et al.
Pubblicazione: (2024)
di: Pi, Renjie, et al.
Pubblicazione: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
di: Wang, Xinyu, et al.
Pubblicazione: (2025)
di: Wang, Xinyu, et al.
Pubblicazione: (2025)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
di: Chen, Liangyu, et al.
Pubblicazione: (2025)
Causal-SAM-LLM: Large Language Models as Causal Reasoners for Robust Medical Segmentation
di: Tang, Tao, et al.
Pubblicazione: (2025)
di: Tang, Tao, et al.
Pubblicazione: (2025)
GenieBlue: Integrating both Linguistic and Multimodal Capabilities for Large Language Models on Mobile Devices
di: Lu, Xudong, et al.
Pubblicazione: (2025)
di: Lu, Xudong, et al.
Pubblicazione: (2025)
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models
di: Chen, Junzhe, et al.
Pubblicazione: (2024)
di: Chen, Junzhe, et al.
Pubblicazione: (2024)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
di: Zhu, Wenxin, et al.
Pubblicazione: (2025)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
di: Chen, Jiaxing, et al.
Pubblicazione: (2024)
di: Chen, Jiaxing, et al.
Pubblicazione: (2024)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
di: Zhan, Xiaoyu, et al.
Pubblicazione: (2025)
di: Zhan, Xiaoyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CELLO: Causal Evaluation of Large Vision-Language Models
di: Chen, Meiqi, et al.
Pubblicazione: (2024) -
Detecting Offensive Memes with Social Biases in Singapore Context Using Multimodal Large Language Models
di: Yuxuan, Cao, et al.
Pubblicazione: (2025) -
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
di: Li, Zhuowan, et al.
Pubblicazione: (2022) -
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
di: Wang, Weihang, et al.
Pubblicazione: (2025) -
Studying and Mitigating Biases in Sign Language Understanding Models
di: Atwell, Katherine, et al.
Pubblicazione: (2024)