Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917881264996352 |
|---|---|
| author | Wang, Xuechen Zhao, Shiwan Sun, Haoqin Wang, Hui Zhou, Jiaming Qin, Yong |
| author_facet | Wang, Xuechen Zhao, Shiwan Sun, Haoqin Wang, Hui Zhou, Jiaming Qin, Yong |
| contents | Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features across these modalities is significant, with most existing approaches adopting a singular alignment strategy. Such a narrow focus not only limits model performance but also fails to address the complexity and ambiguity inherent in emotional expressions. In response, this paper introduces a Multi-Granularity Cross-Modal Alignment (MGCMA) framework, distinguished by its comprehensive approach encompassing distribution-based, instance-based, and token-based alignment modules. This framework enables a multi-level perception of emotional information across modalities. Our experiments on IEMOCAP demonstrate that our proposed method outperforms current state-of-the-art techniques. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_20821 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment Wang, Xuechen Zhao, Shiwan Sun, Haoqin Wang, Hui Zhou, Jiaming Qin, Yong Audio and Speech Processing Computation and Language Sound Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of aligning features across these modalities is significant, with most existing approaches adopting a singular alignment strategy. Such a narrow focus not only limits model performance but also fails to address the complexity and ambiguity inherent in emotional expressions. In response, this paper introduces a Multi-Granularity Cross-Modal Alignment (MGCMA) framework, distinguished by its comprehensive approach encompassing distribution-based, instance-based, and token-based alignment modules. This framework enables a multi-level perception of emotional information across modalities. Our experiments on IEMOCAP demonstrate that our proposed method outperforms current state-of-the-art techniques. |
| title | Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment |
| topic | Audio and Speech Processing Computation and Language Sound |
| url | https://arxiv.org/abs/2412.20821 |