Calibrated Multimodal Representation Learning with Missing Modalities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xiaohao, Xia, Xiaobo, Wei, Jiaheng, Yang, Shuo, Su, Xiu, Ng, See-Kiong, Chua, Tat-Seng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916002461122560
author Liu, Xiaohao
Xia, Xiaobo
Wei, Jiaheng
Yang, Shuo
Su, Xiu
Ng, See-Kiong
Chua, Tat-Seng
author_facet Liu, Xiaohao
Xia, Xiaobo
Wei, Jiaheng
Yang, Shuo
Su, Xiu
Ng, See-Kiong
Chua, Tat-Seng
contents Multimodal representation learning harmonizes distinct modalities by aligning them into a unified latent space. Recent research generalizes traditional cross-modal alignment to produce enhanced multimodal synergy but requires all modalities to be present for a common instance, making it challenging to utilize prevalent datasets with missing modalities. We provide theoretical insights into this issue from an anchor shift perspective. Observed modalities are aligned with a local anchor that deviates from the optimal one when all modalities are present, resulting in an inevitable shift. To address this, we propose CalMRL to calibrate incomplete alignments caused by missing modalities. CalMRL leverages the priors and the inherent connections among modalities to model the imputation for the missing ones at the representation level. To resolve the optimization dilemma, we employ a bi-step learning method with the closed-form solution of the posterior distribution of shared latents. We validate its mitigation of anchor shift and convergence with theoretical guidance. By equipping the calibrated alignment with the existing advanced method, we offer new flexibility to absorb data with missing modalities, which is originally unattainable. Extensive experiments demonstrate the superiority of CalMRL. The code is released at https://github.com/Xiaohao-Liu/CalMRL.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12034
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Calibrated Multimodal Representation Learning with Missing Modalities
Liu, Xiaohao
Xia, Xiaobo
Wei, Jiaheng
Yang, Shuo
Su, Xiu
Ng, See-Kiong
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Machine Learning
Multimedia
Multimodal representation learning harmonizes distinct modalities by aligning them into a unified latent space. Recent research generalizes traditional cross-modal alignment to produce enhanced multimodal synergy but requires all modalities to be present for a common instance, making it challenging to utilize prevalent datasets with missing modalities. We provide theoretical insights into this issue from an anchor shift perspective. Observed modalities are aligned with a local anchor that deviates from the optimal one when all modalities are present, resulting in an inevitable shift. To address this, we propose CalMRL to calibrate incomplete alignments caused by missing modalities. CalMRL leverages the priors and the inherent connections among modalities to model the imputation for the missing ones at the representation level. To resolve the optimization dilemma, we employ a bi-step learning method with the closed-form solution of the posterior distribution of shared latents. We validate its mitigation of anchor shift and convergence with theoretical guidance. By equipping the calibrated alignment with the existing advanced method, we offer new flexibility to absorb data with missing modalities, which is originally unattainable. Extensive experiments demonstrate the superiority of CalMRL. The code is released at https://github.com/Xiaohao-Liu/CalMRL.
title Calibrated Multimodal Representation Learning with Missing Modalities
topic Computer Vision and Pattern Recognition
Machine Learning
Multimedia
url https://arxiv.org/abs/2511.12034