Decoupling Common and Unique Representations for Multimodal Self-supervised Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yi, Albrecht, Conrad M, Braham, Nassim Ait Ali, Liu, Chenying, Xiong, Zhitong, Zhu, Xiao Xiang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917727313068032
author Wang, Yi
Albrecht, Conrad M
Braham, Nassim Ait Ali
Liu, Chenying
Xiong, Zhitong
Zhu, Xiao Xiang
author_facet Wang, Yi
Albrecht, Conrad M
Braham, Nassim Ait Ali
Liu, Chenying
Xiong, Zhitong
Zhu, Xiao Xiang
contents The increasing availability of multi-sensor data sparks wide interest in multimodal self-supervised learning. However, most existing approaches learn only common representations across modalities while ignoring intra-modal training and modality-unique representations. We propose Decoupling Common and Unique Representations (DeCUR), a simple yet effective method for multimodal self-supervised learning. By distinguishing inter- and intra-modal embeddings through multimodal redundancy reduction, DeCUR can integrate complementary information across different modalities. We evaluate DeCUR in three common multimodal scenarios (radar-optical, RGB-elevation, and RGB-depth), and demonstrate its consistent improvement regardless of architectures and for both multimodal and modality-missing settings. With thorough experiments and comprehensive analysis, we hope this work can provide valuable insights and raise more interest in researching the hidden relationships of multimodal representations.
format Preprint
id arxiv_https___arxiv_org_abs_2309_05300
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Decoupling Common and Unique Representations for Multimodal Self-supervised Learning
Wang, Yi
Albrecht, Conrad M
Braham, Nassim Ait Ali
Liu, Chenying
Xiong, Zhitong
Zhu, Xiao Xiang
Computer Vision and Pattern Recognition
The increasing availability of multi-sensor data sparks wide interest in multimodal self-supervised learning. However, most existing approaches learn only common representations across modalities while ignoring intra-modal training and modality-unique representations. We propose Decoupling Common and Unique Representations (DeCUR), a simple yet effective method for multimodal self-supervised learning. By distinguishing inter- and intra-modal embeddings through multimodal redundancy reduction, DeCUR can integrate complementary information across different modalities. We evaluate DeCUR in three common multimodal scenarios (radar-optical, RGB-elevation, and RGB-depth), and demonstrate its consistent improvement regardless of architectures and for both multimodal and modality-missing settings. With thorough experiments and comprehensive analysis, we hope this work can provide valuable insights and raise more interest in researching the hidden relationships of multimodal representations.
title Decoupling Common and Unique Representations for Multimodal Self-supervised Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2309.05300