MMP: Towards Robust Multi-Modal Learning with Masked Modality Projection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nezakati, Niki, Reza, Md Kaykobad, Patil, Ameya, Solh, Mashhour, Asif, M. Salman
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929530889830400
author Nezakati, Niki
Reza, Md Kaykobad
Patil, Ameya
Solh, Mashhour
Asif, M. Salman
author_facet Nezakati, Niki
Reza, Md Kaykobad
Patil, Ameya
Solh, Mashhour
Asif, M. Salman
contents Multimodal learning seeks to combine data from multiple input sources to enhance the performance of different downstream tasks. In real-world scenarios, performance can degrade substantially if some input modalities are missing. Existing methods that can handle missing modalities involve custom training or adaptation steps for each input modality combination. These approaches are either tied to specific modalities or become computationally expensive as the number of input modalities increases. In this paper, we propose Masked Modality Projection (MMP), a method designed to train a single model that is robust to any missing modality scenario. We achieve this by randomly masking a subset of modalities during training and learning to project available input modalities to estimate the tokens for the masked modalities. This approach enables the model to effectively learn to leverage the information from the available modalities to compensate for the missing ones, enhancing missing modality robustness. We conduct a series of experiments with various baseline models and datasets to assess the effectiveness of this strategy. Experiments demonstrate that our approach improves robustness to different missing modality scenarios, outperforming existing methods designed for missing modalities or specific modality combinations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03010
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MMP: Towards Robust Multi-Modal Learning with Masked Modality Projection
Nezakati, Niki
Reza, Md Kaykobad
Patil, Ameya
Solh, Mashhour
Asif, M. Salman
Machine Learning
Computer Vision and Pattern Recognition
Multimodal learning seeks to combine data from multiple input sources to enhance the performance of different downstream tasks. In real-world scenarios, performance can degrade substantially if some input modalities are missing. Existing methods that can handle missing modalities involve custom training or adaptation steps for each input modality combination. These approaches are either tied to specific modalities or become computationally expensive as the number of input modalities increases. In this paper, we propose Masked Modality Projection (MMP), a method designed to train a single model that is robust to any missing modality scenario. We achieve this by randomly masking a subset of modalities during training and learning to project available input modalities to estimate the tokens for the masked modalities. This approach enables the model to effectively learn to leverage the information from the available modalities to compensate for the missing ones, enhancing missing modality robustness. We conduct a series of experiments with various baseline models and datasets to assess the effectiveness of this strategy. Experiments demonstrate that our approach improves robustness to different missing modality scenarios, outperforming existing methods designed for missing modalities or specific modality combinations.
title MMP: Towards Robust Multi-Modal Learning with Masked Modality Projection
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.03010