Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bao, Jindi, Qian, Jianjun, Yan, Mengkai, Yang, Jian
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911501115195392
author Bao, Jindi
Qian, Jianjun
Yan, Mengkai
Yang, Jian
author_facet Bao, Jindi
Qian, Jianjun
Yan, Mengkai
Yang, Jian
contents Multimodal Sentiment Analysis (MSA) seeks to infer human emotions by integrating textual, acoustic, and visual cues. However, existing approaches often rely on all modalities are completeness, whereas real-world applications frequently encounter noise, hardware failures, or privacy restrictions that result in missing modalities. There exists a significant feature misalignment between incomplete and complete modalities, and directly fusing them may even distort the well-learned representations of the intact modalities. To this end, we propose PRLF, a Progressive Representation Learning Framework designed for MSA under uncertain missing-modality conditions. PRLF introduces an Adaptive Modality Reliability Estimator (AMRE), which dynamically quantifies the reliability of each modality using recognition confidence and Fisher information to determine the dominant modality. In addition, the Progressive Interaction (ProgInteract) module iteratively aligns the other modalities with the dominant one, thereby enhancing cross-modal consistency while suppressing noise. Extensive experiments on CMU-MOSI, CMU-MOSEI, and SIMS verify that PRLF outperforms state-of-the-art methods across both inter- and intra-modality missing scenarios, demonstrating its robustness and generalization capability.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09111
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities
Bao, Jindi
Qian, Jianjun
Yan, Mengkai
Yang, Jian
Computer Vision and Pattern Recognition
Multimodal Sentiment Analysis (MSA) seeks to infer human emotions by integrating textual, acoustic, and visual cues. However, existing approaches often rely on all modalities are completeness, whereas real-world applications frequently encounter noise, hardware failures, or privacy restrictions that result in missing modalities. There exists a significant feature misalignment between incomplete and complete modalities, and directly fusing them may even distort the well-learned representations of the intact modalities. To this end, we propose PRLF, a Progressive Representation Learning Framework designed for MSA under uncertain missing-modality conditions. PRLF introduces an Adaptive Modality Reliability Estimator (AMRE), which dynamically quantifies the reliability of each modality using recognition confidence and Fisher information to determine the dominant modality. In addition, the Progressive Interaction (ProgInteract) module iteratively aligns the other modalities with the dominant one, thereby enhancing cross-modal consistency while suppressing noise. Extensive experiments on CMU-MOSI, CMU-MOSEI, and SIMS verify that PRLF outperforms state-of-the-art methods across both inter- and intra-modality missing scenarios, demonstrating its robustness and generalization capability.
title Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.09111