MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866929399714021376 |
|---|---|
| author | Sun, Haiyang Zhang, Fulin Gao, Yingying Lian, Zheng Zhang, Shilei Feng, Junlan |
| author_facet | Sun, Haiyang Zhang, Fulin Gao, Yingying Lian, Zheng Zhang, Shilei Feng, Junlan |
| contents | Speech Emotion Recognition (SER) is an important research topic in human-computer interaction. Many recent works focus on directly extracting emotional cues through pre-trained knowledge, frequently overlooking considerations of appropriateness and comprehensiveness. Therefore, we propose a novel framework for pre-training knowledge in SER, called Multi-perspective Fusion Search Network (MFSN). Considering comprehensiveness, we partition speech knowledge into Textual-related Emotional Content (TEC) and Speech-related Emotional Content (SEC), capturing cues from both semantic and acoustic perspectives, and we design a new architecture search space to fully leverage them. Considering appropriateness, we verify the efficacy of different modeling approaches in capturing SEC and fills the gap in current research. Experimental results on multiple datasets demonstrate the superiority of MFSN. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2306_09361 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition Sun, Haiyang Zhang, Fulin Gao, Yingying Lian, Zheng Zhang, Shilei Feng, Junlan Audio and Speech Processing Computation and Language Sound Speech Emotion Recognition (SER) is an important research topic in human-computer interaction. Many recent works focus on directly extracting emotional cues through pre-trained knowledge, frequently overlooking considerations of appropriateness and comprehensiveness. Therefore, we propose a novel framework for pre-training knowledge in SER, called Multi-perspective Fusion Search Network (MFSN). Considering comprehensiveness, we partition speech knowledge into Textual-related Emotional Content (TEC) and Speech-related Emotional Content (SEC), capturing cues from both semantic and acoustic perspectives, and we design a new architecture search space to fully leverage them. Considering appropriateness, we verify the efficacy of different modeling approaches in capturing SEC and fills the gap in current research. Experimental results on multiple datasets demonstrate the superiority of MFSN. |
| title | MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition |
| topic | Audio and Speech Processing Computation and Language Sound |
| url | https://arxiv.org/abs/2306.09361 |