MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Haiyang, Zhang, Fulin, Gao, Yingying, Lian, Zheng, Zhang, Shilei, Feng, Junlan
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929399714021376
author Sun, Haiyang
Zhang, Fulin
Gao, Yingying
Lian, Zheng
Zhang, Shilei
Feng, Junlan
author_facet Sun, Haiyang
Zhang, Fulin
Gao, Yingying
Lian, Zheng
Zhang, Shilei
Feng, Junlan
contents Speech Emotion Recognition (SER) is an important research topic in human-computer interaction. Many recent works focus on directly extracting emotional cues through pre-trained knowledge, frequently overlooking considerations of appropriateness and comprehensiveness. Therefore, we propose a novel framework for pre-training knowledge in SER, called Multi-perspective Fusion Search Network (MFSN). Considering comprehensiveness, we partition speech knowledge into Textual-related Emotional Content (TEC) and Speech-related Emotional Content (SEC), capturing cues from both semantic and acoustic perspectives, and we design a new architecture search space to fully leverage them. Considering appropriateness, we verify the efficacy of different modeling approaches in capturing SEC and fills the gap in current research. Experimental results on multiple datasets demonstrate the superiority of MFSN.
format Preprint
id arxiv_https___arxiv_org_abs_2306_09361
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
Sun, Haiyang
Zhang, Fulin
Gao, Yingying
Lian, Zheng
Zhang, Shilei
Feng, Junlan
Audio and Speech Processing
Computation and Language
Sound
Speech Emotion Recognition (SER) is an important research topic in human-computer interaction. Many recent works focus on directly extracting emotional cues through pre-trained knowledge, frequently overlooking considerations of appropriateness and comprehensiveness. Therefore, we propose a novel framework for pre-training knowledge in SER, called Multi-perspective Fusion Search Network (MFSN). Considering comprehensiveness, we partition speech knowledge into Textual-related Emotional Content (TEC) and Speech-related Emotional Content (SEC), capturing cues from both semantic and acoustic perspectives, and we design a new architecture search space to fully leverage them. Considering appropriateness, we verify the efficacy of different modeling approaches in capturing SEC and fills the gap in current research. Experimental results on multiple datasets demonstrate the superiority of MFSN.
title MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2306.09361