MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Shuchang, Naseem, Usman, Meng, Mingyuan, Feng, Dagan, Kim, Jinman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913854534975488
author Ye, Shuchang
Naseem, Usman
Meng, Mingyuan
Feng, Dagan
Kim, Jinman
author_facet Ye, Shuchang
Naseem, Usman
Meng, Mingyuan
Feng, Dagan
Kim, Jinman
contents Medical Visual Question Answering (MedVQA) is crucial for enhancing the efficiency of clinical diagnosis by providing accurate and timely responses to clinicians' inquiries regarding medical images. Existing MedVQA models suffered from modality preference bias, where predictions are heavily dominated by one modality while overlooking the other (in MedVQA, usually questions dominate the answer but images are overlooked), thereby failing to learn multimodal knowledge. To overcome the modality preference bias, we proposed a Medical CounterFactual VQA (MedCFVQA) model, which trains with bias and leverages causal graphs to eliminate the modality preference bias during inference. Existing MedVQA datasets exhibit substantial prior dependencies between questions and answers, which results in acceptable performance even if the model significantly suffers from the modality preference bias. To address this issue, we reconstructed new datasets by leveraging existing MedVQA datasets and Changed their P3rior dependencies (CP) between questions and their answers in the training and test set. Extensive experiments demonstrate that MedCFVQA significantly outperforms its non-causal counterpart on both SLAKE, RadVQA and SLAKE-CP, RadVQA-CP datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16209
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
Ye, Shuchang
Naseem, Usman
Meng, Mingyuan
Feng, Dagan
Kim, Jinman
Computer Vision and Pattern Recognition
Medical Visual Question Answering (MedVQA) is crucial for enhancing the efficiency of clinical diagnosis by providing accurate and timely responses to clinicians' inquiries regarding medical images. Existing MedVQA models suffered from modality preference bias, where predictions are heavily dominated by one modality while overlooking the other (in MedVQA, usually questions dominate the answer but images are overlooked), thereby failing to learn multimodal knowledge. To overcome the modality preference bias, we proposed a Medical CounterFactual VQA (MedCFVQA) model, which trains with bias and leverages causal graphs to eliminate the modality preference bias during inference. Existing MedVQA datasets exhibit substantial prior dependencies between questions and answers, which results in acceptable performance even if the model significantly suffers from the modality preference bias. To address this issue, we reconstructed new datasets by leveraging existing MedVQA datasets and Changed their P3rior dependencies (CP) between questions and their answers in the training and test set. Extensive experiments demonstrate that MedCFVQA significantly outperforms its non-causal counterpart on both SLAKE, RadVQA and SLAKE-CP, RadVQA-CP datasets.
title MedCFVQA: A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16209