Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Jinming, Fang, Jingyi, Zheng, Yuanzhong, Wang, Yaoxuan, Fei, Haojun
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913765527650304
author Chen, Jinming
Fang, Jingyi
Zheng, Yuanzhong
Wang, Yaoxuan
Fei, Haojun
author_facet Chen, Jinming
Fang, Jingyi
Zheng, Yuanzhong
Wang, Yaoxuan
Fei, Haojun
contents Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving the recognition accuracy. However, the lack of high-quality multimodal data and the challenge of achieving optimal alignment between different modalities significantly limit the potential for improvement in multimodal approaches. In this paper, the proposed Qieemo framework effectively utilizes the pretrained automatic speech recognition (ASR) model backbone which contains naturally frame aligned textual and emotional features, to achieve precise emotion classification solely based on the audio modality. Furthermore, we design the multimodal fusion (MMF) module and cross-modal attention (CMA) module in order to fuse the phonetic posteriorgram (PPG) and emotional features extracted by the ASR encoder for improving recognition accuracy. The experimental results on the IEMOCAP dataset demonstrate that Qieemo outperforms the benchmark unimodal, multimodal, and self-supervised models with absolute improvements of 3.0%, 1.2%, and 1.9% respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2503_22687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations
Chen, Jinming
Fang, Jingyi
Zheng, Yuanzhong
Wang, Yaoxuan
Fei, Haojun
Audio and Speech Processing
Artificial Intelligence
Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving the recognition accuracy. However, the lack of high-quality multimodal data and the challenge of achieving optimal alignment between different modalities significantly limit the potential for improvement in multimodal approaches. In this paper, the proposed Qieemo framework effectively utilizes the pretrained automatic speech recognition (ASR) model backbone which contains naturally frame aligned textual and emotional features, to achieve precise emotion classification solely based on the audio modality. Furthermore, we design the multimodal fusion (MMF) module and cross-modal attention (CMA) module in order to fuse the phonetic posteriorgram (PPG) and emotional features extracted by the ASR encoder for improving recognition accuracy. The experimental results on the IEMOCAP dataset demonstrate that Qieemo outperforms the benchmark unimodal, multimodal, and self-supervised models with absolute improvements of 3.0%, 1.2%, and 1.9% respectively.
title Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations
topic Audio and Speech Processing
Artificial Intelligence
url https://arxiv.org/abs/2503.22687