Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Xuechen, Zhao, Shiwan, Sun, Haoqin, Wang, Hui, Zhou, Jiaming, Qin, Yong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917881264996352
author Wang, Xuechen
Zhao, Shiwan
Sun, Haoqin
Wang, Hui
Zhou, Jiaming
Qin, Yong
author_facet Wang, Xuechen
Zhao, Shiwan
Sun, Haoqin
Wang, Hui
Zhou, Jiaming
Qin, Yong
contents Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features across these modalities is significant, with most existing approaches adopting a singular alignment strategy. Such a narrow focus not only limits model performance but also fails to address the complexity and ambiguity inherent in emotional expressions. In response, this paper introduces a Multi-Granularity Cross-Modal Alignment (MGCMA) framework, distinguished by its comprehensive approach encompassing distribution-based, instance-based, and token-based alignment modules. This framework enables a multi-level perception of emotional information across modalities. Our experiments on IEMOCAP demonstrate that our proposed method outperforms current state-of-the-art techniques.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20821
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
Wang, Xuechen
Zhao, Shiwan
Sun, Haoqin
Wang, Hui
Zhou, Jiaming
Qin, Yong
Audio and Speech Processing
Computation and Language
Sound
Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features across these modalities is significant, with most existing approaches adopting a singular alignment strategy. Such a narrow focus not only limits model performance but also fails to address the complexity and ambiguity inherent in emotional expressions. In response, this paper introduces a Multi-Granularity Cross-Modal Alignment (MGCMA) framework, distinguished by its comprehensive approach encompassing distribution-based, instance-based, and token-based alignment modules. This framework enables a multi-level perception of emotional information across modalities. Our experiments on IEMOCAP demonstrate that our proposed method outperforms current state-of-the-art techniques.
title Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2412.20821