TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Quan, Weize, Feng, Yunfei, Zhou, Ming, Zhao, Yunzhen, Wang, Tong, Yan, Dong-Ming
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910931939753984
author Quan, Weize
Feng, Yunfei
Zhou, Ming
Zhao, Yunzhen
Wang, Tong
Yan, Dong-Ming
author_facet Quan, Weize
Feng, Yunfei
Zhou, Ming
Zhao, Yunzhen
Wang, Tong
Yan, Dong-Ming
contents Multimodal Sentiment Analysis (MSA) endeavors to understand human sentiment by leveraging language, visual, and acoustic modalities. Despite the remarkable performance exhibited by previous MSA approaches, the presence of inherent multimodal heterogeneities poses a challenge, with the contribution of different modalities varying considerably. Past research predominantly focused on improving representation learning techniques and feature fusion strategies. However, many of these efforts overlooked the variation in semantic richness among different modalities, treating each modality uniformly. This approach may lead to underestimating the significance of strong modalities while overemphasizing the importance of weak ones. Motivated by these insights, we introduce a Text-oriented Cross-Attention Network (TCAN), emphasizing the predominant role of the text modality in MSA. Specifically, for each multimodal sample, by taking unaligned sequences of the three modalities as inputs, we initially allocate the extracted unimodal features into a visual-text and an acoustic-text pair. Subsequently, we implement self-attention on the text modality and apply text-queried cross-attention to the visual and acoustic modalities. To mitigate the influence of noise signals and redundant features, we incorporate a gated control mechanism into the framework. Additionally, we introduce unimodal joint learning to gain a deeper understanding of homogeneous emotional tendencies across diverse modalities through backpropagation. Experimental results demonstrate that TCAN consistently outperforms state-of-the-art MSA methods on two datasets (CMU-MOSI and CMU-MOSEI).
format Preprint
id arxiv_https___arxiv_org_abs_2404_04545
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis
Quan, Weize
Feng, Yunfei
Zhou, Ming
Zhao, Yunzhen
Wang, Tong
Yan, Dong-Ming
Multimedia
Computation and Language
Multimodal Sentiment Analysis (MSA) endeavors to understand human sentiment by leveraging language, visual, and acoustic modalities. Despite the remarkable performance exhibited by previous MSA approaches, the presence of inherent multimodal heterogeneities poses a challenge, with the contribution of different modalities varying considerably. Past research predominantly focused on improving representation learning techniques and feature fusion strategies. However, many of these efforts overlooked the variation in semantic richness among different modalities, treating each modality uniformly. This approach may lead to underestimating the significance of strong modalities while overemphasizing the importance of weak ones. Motivated by these insights, we introduce a Text-oriented Cross-Attention Network (TCAN), emphasizing the predominant role of the text modality in MSA. Specifically, for each multimodal sample, by taking unaligned sequences of the three modalities as inputs, we initially allocate the extracted unimodal features into a visual-text and an acoustic-text pair. Subsequently, we implement self-attention on the text modality and apply text-queried cross-attention to the visual and acoustic modalities. To mitigate the influence of noise signals and redundant features, we incorporate a gated control mechanism into the framework. Additionally, we introduce unimodal joint learning to gain a deeper understanding of homogeneous emotional tendencies across diverse modalities through backpropagation. Experimental results demonstrate that TCAN consistently outperforms state-of-the-art MSA methods on two datasets (CMU-MOSI and CMU-MOSEI).
title TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis
topic Multimedia
Computation and Language
url https://arxiv.org/abs/2404.04545