See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kwon, JuneHyoung, Kim, MiHyeon, Lee, Eunju, Choi, Juhwan, Kim, YoungBin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909540923998208
author Kwon, JuneHyoung
Kim, MiHyeon
Lee, Eunju
Choi, Juhwan
Kim, YoungBin
author_facet Kwon, JuneHyoung
Kim, MiHyeon
Lee, Eunju
Choi, Juhwan
Kim, YoungBin
contents Vision-language (VL) models have demonstrated strong performance across various tasks. However, these models often rely on a specific modality for predictions, leading to "dominant modality bias.'' This bias significantly hurts performance, especially when one modality is impaired. In this study, we analyze model behavior under dominant modality bias and theoretically show that unaligned gradients or differences in gradient magnitudes prevent balanced convergence of the loss. Based on these findings, we propose a novel framework, BalGrad to mitigate dominant modality bias. Our approach includes inter-modality gradient reweighting, adjusting the gradient of KL divergence based on each modality's contribution, and inter-task gradient projection to align task directions in a non-conflicting manner. Experiments on UPMC Food-101, Hateful Memes, and MM-IMDb datasets confirm that BalGrad effectively alleviates over-reliance on specific modalities when making predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13834
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias
Kwon, JuneHyoung
Kim, MiHyeon
Lee, Eunju
Choi, Juhwan
Kim, YoungBin
Computer Vision and Pattern Recognition
Vision-language (VL) models have demonstrated strong performance across various tasks. However, these models often rely on a specific modality for predictions, leading to "dominant modality bias.'' This bias significantly hurts performance, especially when one modality is impaired. In this study, we analyze model behavior under dominant modality bias and theoretically show that unaligned gradients or differences in gradient magnitudes prevent balanced convergence of the loss. Based on these findings, we propose a novel framework, BalGrad to mitigate dominant modality bias. Our approach includes inter-modality gradient reweighting, adjusting the gradient of KL divergence based on each modality's contribution, and inter-task gradient projection to align task directions in a non-conflicting manner. Experiments on UPMC Food-101, Hateful Memes, and MM-IMDb datasets confirm that BalGrad effectively alleviates over-reliance on specific modalities when making predictions.
title See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.13834