MIDAS: Misalignment-based Data Augmentation Strategy for Imbalanced Multimodal Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hwang, Seong-Hyeon, Choi, Soyoung, Whang, Steven Euijong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914067428409344
author Hwang, Seong-Hyeon
Choi, Soyoung
Whang, Steven Euijong
author_facet Hwang, Seong-Hyeon
Choi, Soyoung
Whang, Steven Euijong
contents Multimodal models often over-rely on dominant modalities, failing to achieve optimal performance. While prior work focuses on modifying training objectives or optimization procedures, data-centric solutions remain underexplored. We propose MIDAS, a novel data augmentation strategy that generates misaligned samples with semantically inconsistent cross-modal information, labeled using unimodal confidence scores to compel learning from contradictory signals. However, this confidence-based labeling can still favor the more confident modality. To address this within our misaligned samples, we introduce weak-modality weighting, which dynamically increases the loss weight of the least confident modality, thereby helping the model fully utilize weaker modality. Furthermore, when misaligned features exhibit greater similarity to the aligned features, these misaligned samples pose a greater challenge, thereby enabling the model to better distinguish between classes. To leverage this, we propose hard-sample weighting, which prioritizes such semantically ambiguous misaligned samples. Experiments on multiple multimodal classification benchmarks demonstrate that MIDAS significantly outperforms related baselines in addressing modality imbalance.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25831
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MIDAS: Misalignment-based Data Augmentation Strategy for Imbalanced Multimodal Learning
Hwang, Seong-Hyeon
Choi, Soyoung
Whang, Steven Euijong
Machine Learning
Multimodal models often over-rely on dominant modalities, failing to achieve optimal performance. While prior work focuses on modifying training objectives or optimization procedures, data-centric solutions remain underexplored. We propose MIDAS, a novel data augmentation strategy that generates misaligned samples with semantically inconsistent cross-modal information, labeled using unimodal confidence scores to compel learning from contradictory signals. However, this confidence-based labeling can still favor the more confident modality. To address this within our misaligned samples, we introduce weak-modality weighting, which dynamically increases the loss weight of the least confident modality, thereby helping the model fully utilize weaker modality. Furthermore, when misaligned features exhibit greater similarity to the aligned features, these misaligned samples pose a greater challenge, thereby enabling the model to better distinguish between classes. To leverage this, we propose hard-sample weighting, which prioritizes such semantically ambiguous misaligned samples. Experiments on multiple multimodal classification benchmarks demonstrate that MIDAS significantly outperforms related baselines in addressing modality imbalance.
title MIDAS: Misalignment-based Data Augmentation Strategy for Imbalanced Multimodal Learning
topic Machine Learning
url https://arxiv.org/abs/2509.25831