AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Sheng, Liu, Jiaxing, Wang, Longbiao, He, Dongxiao, Wang, Xiaobao, Dang, Jianwu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910508515328000
author Wu, Sheng
Liu, Jiaxing
Wang, Longbiao
He, Dongxiao
Wang, Xiaobao
Dang, Jianwu
author_facet Wu, Sheng
Liu, Jiaxing
Wang, Longbiao
He, Dongxiao
Wang, Xiaobao
Dang, Jianwu
contents Emotion Recognition in Conversations (ERC) is a popular task in natural language processing, which aims to recognize the emotional state of the speaker in conversations. While current research primarily emphasizes contextual modeling, there exists a dearth of investigation into effective multimodal fusion methods. We propose a novel framework called AIMDiT to solve the problem of multimodal fusion of deep features. Specifically, we design a Modality Augmentation Network which performs rich representation learning through dimension transformation of different modalities and parameter-efficient inception block. On the other hand, the Modality Interaction Network performs interaction fusion of extracted inter-modal features and intra-modal features. Experiments conducted using our AIMDiT framework on the public benchmark dataset MELD reveal 2.34% and 2.87% improvements in terms of the Acc-7 and w-F1 metrics compared to the state-of-the-art (SOTA) models.
format Preprint
id arxiv_https___arxiv_org_abs_2407_00743
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations
Wu, Sheng
Liu, Jiaxing
Wang, Longbiao
He, Dongxiao
Wang, Xiaobao
Dang, Jianwu
Multimedia
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Emotion Recognition in Conversations (ERC) is a popular task in natural language processing, which aims to recognize the emotional state of the speaker in conversations. While current research primarily emphasizes contextual modeling, there exists a dearth of investigation into effective multimodal fusion methods. We propose a novel framework called AIMDiT to solve the problem of multimodal fusion of deep features. Specifically, we design a Modality Augmentation Network which performs rich representation learning through dimension transformation of different modalities and parameter-efficient inception block. On the other hand, the Modality Interaction Network performs interaction fusion of extracted inter-modal features and intra-modal features. Experiments conducted using our AIMDiT framework on the public benchmark dataset MELD reveal 2.34% and 2.87% improvements in terms of the Acc-7 and w-F1 metrics compared to the state-of-the-art (SOTA) models.
title AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations
topic Multimedia
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2407.00743