A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dai, Wei, Zheng, Dequan, Yu, Feng, Zhang, Yanrong, Hou, Yaohui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909489866735616
author Dai, Wei
Zheng, Dequan
Yu, Feng
Zhang, Yanrong
Hou, Yaohui
author_facet Dai, Wei
Zheng, Dequan
Yu, Feng
Zhang, Yanrong
Hou, Yaohui
contents With the advancement of artificial intelligence and computer vision technologies, multimodal emotion recognition has become a prominent research topic. However, existing methods face challenges such as heterogeneous data fusion and the effective utilization of modality correlations. This paper proposes a novel multimodal emotion recognition approach, DeepMSI-MER, based on the integration of contrastive learning and visual sequence compression. The proposed method enhances cross-modal feature fusion through contrastive learning and reduces redundancy in the visual modality by leveraging visual sequence compression. Experimental results on two public datasets, IEMOCAP and MELD, demonstrate that DeepMSI-MER significantly improves the accuracy and robustness of emotion recognition, validating the effectiveness of multimodal feature fusion and the proposed approach.
format Preprint
id arxiv_https___arxiv_org_abs_2502_08573
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion
Dai, Wei
Zheng, Dequan
Yu, Feng
Zhang, Yanrong
Hou, Yaohui
Computer Vision and Pattern Recognition
Artificial Intelligence
With the advancement of artificial intelligence and computer vision technologies, multimodal emotion recognition has become a prominent research topic. However, existing methods face challenges such as heterogeneous data fusion and the effective utilization of modality correlations. This paper proposes a novel multimodal emotion recognition approach, DeepMSI-MER, based on the integration of contrastive learning and visual sequence compression. The proposed method enhances cross-modal feature fusion through contrastive learning and reduces redundancy in the visual modality by leveraging visual sequence compression. Experimental results on two public datasets, IEMOCAP and MELD, demonstrate that DeepMSI-MER significantly improves the accuracy and robustness of emotion recognition, validating the effectiveness of multimodal feature fusion and the proposed approach.
title A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2502.08573