Mixture of Disentangled Experts with Missing Modalities for Robust Multimodal Sentiment Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xiang, Zhang, Xiaoming, Miao, Dezhuang, Cheng, Xianfu, Li, Dawei, Han, Honggui, Li, Zhoujun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915768577294336
author Li, Xiang
Zhang, Xiaoming
Miao, Dezhuang
Cheng, Xianfu
Li, Dawei
Han, Honggui
Li, Zhoujun
author_facet Li, Xiang
Zhang, Xiaoming
Miao, Dezhuang
Cheng, Xianfu
Li, Dawei
Han, Honggui
Li, Zhoujun
contents Multimodal Sentiment Analysis (MSA) integrates multiple modalities to infer human sentiment, but real-world noise often leads to missing or corrupted data. However, existing feature-disentangled methods struggle to handle the internal variations of heterogeneous information under uncertain missingness, making it difficult to learn effective multimodal representations from degraded modalities. To address this issue, we propose DERL, a Disentangled Expert Representation Learning framework for robust MSA. Specifically, DERL employs hybrid experts to adaptively disentangle multimodal inputs into orthogonal private and shared representation spaces. A multi-level reconstruction strategy is further developed to provide collaborative supervision, enhancing both the expressiveness and robustness of the learned representations. Finally, the disentangled features act as modality experts with distinct roles to generate importance-aware fusion results. Extensive experiments on two MSA benchmarks demonstrate that DERL outperforms state-of-the-art methods under various missing-modality conditions. For instance, our method achieves improvements of 2.47% in Acc-2 and 2.25% in MAE on MOSI under intra-modal missingness.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01833
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mixture of Disentangled Experts with Missing Modalities for Robust Multimodal Sentiment Analysis
Li, Xiang
Zhang, Xiaoming
Miao, Dezhuang
Cheng, Xianfu
Li, Dawei
Han, Honggui
Li, Zhoujun
Multimedia
Multimodal Sentiment Analysis (MSA) integrates multiple modalities to infer human sentiment, but real-world noise often leads to missing or corrupted data. However, existing feature-disentangled methods struggle to handle the internal variations of heterogeneous information under uncertain missingness, making it difficult to learn effective multimodal representations from degraded modalities. To address this issue, we propose DERL, a Disentangled Expert Representation Learning framework for robust MSA. Specifically, DERL employs hybrid experts to adaptively disentangle multimodal inputs into orthogonal private and shared representation spaces. A multi-level reconstruction strategy is further developed to provide collaborative supervision, enhancing both the expressiveness and robustness of the learned representations. Finally, the disentangled features act as modality experts with distinct roles to generate importance-aware fusion results. Extensive experiments on two MSA benchmarks demonstrate that DERL outperforms state-of-the-art methods under various missing-modality conditions. For instance, our method achieves improvements of 2.47% in Acc-2 and 2.25% in MAE on MOSI under intra-modal missingness.
title Mixture of Disentangled Experts with Missing Modalities for Robust Multimodal Sentiment Analysis
topic Multimedia
url https://arxiv.org/abs/2602.01833