Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Chunlei, Zhou, Ziyang, He, Lucas, Du, Xiaojing, Ouyang, Chun, Gan, Zhongxue
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918296732827648
author Meng, Chunlei
Zhou, Ziyang
He, Lucas
Du, Xiaojing
Ouyang, Chun
Gan, Zhongxue
author_facet Meng, Chunlei
Zhou, Ziyang
He, Lucas
Du, Xiaojing
Ouyang, Chun
Gan, Zhongxue
contents Multimodal Sentiment Analysis integrates Linguistic, Visual, and Acoustic. Mainstream approaches based on modality-invariant and modality-specific factorization or on complex fusion still rely on spatiotemporal mixed modeling. This ignores spatiotemporal heterogeneity, leading to spatiotemporal information asymmetry and thus limited performance. Hence, we propose TSDA, Temporal-Spatial Decouple before Act, which explicitly decouples each modality into temporal dynamics and spatial structural context before any interaction. For every modality, a temporal encoder and a spatial encoder project signals into separate temporal and spatial body. Factor-Consistent Cross-Modal Alignment then aligns temporal features only with their temporal counterparts across modalities, and spatial features only with their spatial counterparts. Factor specific supervision and decorrelation regularization reduce cross factor leakage while preserving complementarity. A Gated Recouple module subsequently recouples the aligned streams for task. Extensive experiments show that TSDA outperforms baselines. Ablation analysis studies confirm the necessity and interpretability of the design.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13659
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
Meng, Chunlei
Zhou, Ziyang
He, Lucas
Du, Xiaojing
Ouyang, Chun
Gan, Zhongxue
Computation and Language
Artificial Intelligence
Multimedia
Multimodal Sentiment Analysis integrates Linguistic, Visual, and Acoustic. Mainstream approaches based on modality-invariant and modality-specific factorization or on complex fusion still rely on spatiotemporal mixed modeling. This ignores spatiotemporal heterogeneity, leading to spatiotemporal information asymmetry and thus limited performance. Hence, we propose TSDA, Temporal-Spatial Decouple before Act, which explicitly decouples each modality into temporal dynamics and spatial structural context before any interaction. For every modality, a temporal encoder and a spatial encoder project signals into separate temporal and spatial body. Factor-Consistent Cross-Modal Alignment then aligns temporal features only with their temporal counterparts across modalities, and spatial features only with their spatial counterparts. Factor specific supervision and decorrelation regularization reduce cross factor leakage while preserving complementarity. A Gated Recouple module subsequently recouples the aligned streams for task. Extensive experiments show that TSDA outperforms baselines. Ablation analysis studies confirm the necessity and interpretability of the design.
title Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
topic Computation and Language
Artificial Intelligence
Multimedia
url https://arxiv.org/abs/2601.13659