Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929445884919808 |
|---|---|
| author | Sun, Haoqin Zhao, Shiwan Kong, Xiangyu Wang, Xuechen Wang, Hui Zhou, Jiaming Qin, Yong |
| author_facet | Sun, Haoqin Zhao, Shiwan Kong, Xiangyu Wang, Xuechen Wang, Hui Zhou, Jiaming Qin, Yong |
| contents | Recognizing emotions from speech is a daunting task due to the subtlety and ambiguity of expressions. Traditional speech emotion recognition (SER) systems, which typically rely on a singular, precise emotion label, struggle with this complexity. Therefore, modeling the inherent ambiguity of emotions is an urgent problem. In this paper, we propose an iterative prototype refinement framework (IPR) for ambiguous SER. IPR comprises two interlinked components: contrastive learning and class prototypes. The former provides an efficient way to obtain high-quality representations of ambiguous samples. The latter are dynamically updated based on ambiguous labels -- the similarity of the ambiguous data to all prototypes. These refined embeddings yield precise pseudo labels, thus reinforcing representation quality. Experimental evaluations conducted on the IEMOCAP dataset validate the superior performance of IPR over state-of-the-art methods, thus proving the effectiveness of our proposed method. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_00325 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition Sun, Haoqin Zhao, Shiwan Kong, Xiangyu Wang, Xuechen Wang, Hui Zhou, Jiaming Qin, Yong Sound Audio and Speech Processing Recognizing emotions from speech is a daunting task due to the subtlety and ambiguity of expressions. Traditional speech emotion recognition (SER) systems, which typically rely on a singular, precise emotion label, struggle with this complexity. Therefore, modeling the inherent ambiguity of emotions is an urgent problem. In this paper, we propose an iterative prototype refinement framework (IPR) for ambiguous SER. IPR comprises two interlinked components: contrastive learning and class prototypes. The former provides an efficient way to obtain high-quality representations of ambiguous samples. The latter are dynamically updated based on ambiguous labels -- the similarity of the ambiguous data to all prototypes. These refined embeddings yield precise pseudo labels, thus reinforcing representation quality. Experimental evaluations conducted on the IEMOCAP dataset validate the superior performance of IPR over state-of-the-art methods, thus proving the effectiveness of our proposed method. |
| title | Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2408.00325 |