Adapting WavLM for Speech Emotion Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866909193914548224 |
|---|---|
| author | Diatlova, Daria Udalov, Anton Shutov, Vitalii Spirin, Egor |
| author_facet | Diatlova, Daria Udalov, Anton Shutov, Vitalii Spirin, Egor |
| contents | Recently, the usage of speech self-supervised models (SSL) for downstream tasks has been drawing a lot of attention. While large pre-trained models commonly outperform smaller models trained from scratch, questions regarding the optimal fine-tuning strategies remain prevalent. In this paper, we explore the fine-tuning strategies of the WavLM Large model for the speech emotion recognition task on the MSP Podcast Corpus. More specifically, we perform a series of experiments focusing on using gender and semantic information from utterances. We then sum up our findings and describe the final model we used for submission to Speech Emotion Recognition Challenge 2024. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_04485 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Adapting WavLM for Speech Emotion Recognition Diatlova, Daria Udalov, Anton Shutov, Vitalii Spirin, Egor Machine Learning Sound Audio and Speech Processing Recently, the usage of speech self-supervised models (SSL) for downstream tasks has been drawing a lot of attention. While large pre-trained models commonly outperform smaller models trained from scratch, questions regarding the optimal fine-tuning strategies remain prevalent. In this paper, we explore the fine-tuning strategies of the WavLM Large model for the speech emotion recognition task on the MSP Podcast Corpus. More specifically, we perform a series of experiments focusing on using gender and semantic information from utterances. We then sum up our findings and describe the final model we used for submission to Speech Emotion Recognition Challenge 2024. |
| title | Adapting WavLM for Speech Emotion Recognition |
| topic | Machine Learning Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2405.04485 |