Exploring Missing Modality in Multimodal Egocentric Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramazanova, Merey, Pardo, Alejandro, Alwassel, Humam, Ghanem, Bernard
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866907869396336640
author Ramazanova, Merey
Pardo, Alejandro
Alwassel, Humam
Ghanem, Bernard
author_facet Ramazanova, Merey
Pardo, Alejandro
Alwassel, Humam
Ghanem, Bernard
contents Multimodal video understanding is crucial for analyzing egocentric videos, where integrating multiple sensory signals significantly enhances action recognition and moment localization. However, practical applications often grapple with incomplete modalities due to factors like privacy concerns, efficiency demands, or hardware malfunctions. Addressing this, our study delves into the impact of missing modalities on egocentric action recognition, particularly within transformer-based models. We introduce a novel concept -Missing Modality Token (MMT)-to maintain performance even when modalities are absent, a strategy that proves effective in the Ego4D, Epic-Kitchens, and Epic-Sounds datasets. Our method mitigates the performance loss, reducing it from its original $\sim 30\%$ drop to only $\sim 10\%$ when half of the test set is modal-incomplete. Through extensive experimentation, we demonstrate the adaptability of MMT to different training scenarios and its superiority in handling missing modalities compared to current methods. Our research contributes a comprehensive analysis and an innovative approach, opening avenues for more resilient multimodal systems in real-world settings.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11470
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Missing Modality in Multimodal Egocentric Datasets
Ramazanova, Merey
Pardo, Alejandro
Alwassel, Humam
Ghanem, Bernard
Computer Vision and Pattern Recognition
Multimodal video understanding is crucial for analyzing egocentric videos, where integrating multiple sensory signals significantly enhances action recognition and moment localization. However, practical applications often grapple with incomplete modalities due to factors like privacy concerns, efficiency demands, or hardware malfunctions. Addressing this, our study delves into the impact of missing modalities on egocentric action recognition, particularly within transformer-based models. We introduce a novel concept -Missing Modality Token (MMT)-to maintain performance even when modalities are absent, a strategy that proves effective in the Ego4D, Epic-Kitchens, and Epic-Sounds datasets. Our method mitigates the performance loss, reducing it from its original $\sim 30\%$ drop to only $\sim 10\%$ when half of the test set is modal-incomplete. Through extensive experimentation, we demonstrate the adaptability of MMT to different training scenarios and its superiority in handling missing modalities compared to current methods. Our research contributes a comprehensive analysis and an innovative approach, opening avenues for more resilient multimodal systems in real-world settings.
title Exploring Missing Modality in Multimodal Egocentric Datasets
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.11470