Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Meng, Wei, Yake, Yin, Jianxiong, Rajan, Deepu, Hu, Di, See, Simon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915061093629952
author Shen, Meng
Wei, Yake
Yin, Jianxiong
Rajan, Deepu
Hu, Di
See, Simon
author_facet Shen, Meng
Wei, Yake
Yin, Jianxiong
Rajan, Deepu
Hu, Di
See, Simon
contents Training multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on sufficient labeled data to train a well-calibrated model that can assess the uncertainty and diversity of unlabeled data. However, when assembling a dataset, labeled data are often scarce initially, leading to a cold-start problem. Additionally, most AL methods seldom address multimodal data, highlighting a research gap in this field. Our research addresses these issues by developing a two-stage method for Multi-Modal Cold-Start Active Learning (MMCSAL). Firstly, we observe the modality gap, a significant distance between the centroids of representations from different modalities, when only using cross-modal pairing information as self-supervision signals. This modality gap affects data selection process, as we calculate both uni-modal and cross-modal distances. To address this, we introduce uni-modal prototypes to bridge the modality gap. Secondly, conventional AL methods often falter in multimodal scenarios where alignment between modalities is overlooked. Therefore, we propose enhancing cross-modal alignment through regularization, thereby improving the quality of selected multimodal data pairs in AL. Finally, our experiments demonstrate MMCSAL's efficacy in selecting multimodal data pairs across three multimodal datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09126
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
Shen, Meng
Wei, Yake
Yin, Jianxiong
Rajan, Deepu
Hu, Di
See, Simon
Multimedia
Artificial Intelligence
Machine Learning
Training multimodal models requires a large amount of labeled data. Active learning (AL) aim to reduce labeling costs. Most AL methods employ warm-start approaches, which rely on sufficient labeled data to train a well-calibrated model that can assess the uncertainty and diversity of unlabeled data. However, when assembling a dataset, labeled data are often scarce initially, leading to a cold-start problem. Additionally, most AL methods seldom address multimodal data, highlighting a research gap in this field. Our research addresses these issues by developing a two-stage method for Multi-Modal Cold-Start Active Learning (MMCSAL). Firstly, we observe the modality gap, a significant distance between the centroids of representations from different modalities, when only using cross-modal pairing information as self-supervision signals. This modality gap affects data selection process, as we calculate both uni-modal and cross-modal distances. To address this, we introduce uni-modal prototypes to bridge the modality gap. Secondly, conventional AL methods often falter in multimodal scenarios where alignment between modalities is overlooked. Therefore, we propose enhancing cross-modal alignment through regularization, thereby improving the quality of selected multimodal data pairs in AL. Finally, our experiments demonstrate MMCSAL's efficacy in selecting multimodal data pairs across three multimodal datasets.
title Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
topic Multimedia
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.09126