Adapting WavLM for Speech Emotion Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Diatlova, Daria, Udalov, Anton, Shutov, Vitalii, Spirin, Egor
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909193914548224
author Diatlova, Daria
Udalov, Anton
Shutov, Vitalii
Spirin, Egor
author_facet Diatlova, Daria
Udalov, Anton
Shutov, Vitalii
Spirin, Egor
contents Recently, the usage of speech self-supervised models (SSL) for downstream tasks has been drawing a lot of attention. While large pre-trained models commonly outperform smaller models trained from scratch, questions regarding the optimal fine-tuning strategies remain prevalent. In this paper, we explore the fine-tuning strategies of the WavLM Large model for the speech emotion recognition task on the MSP Podcast Corpus. More specifically, we perform a series of experiments focusing on using gender and semantic information from utterances. We then sum up our findings and describe the final model we used for submission to Speech Emotion Recognition Challenge 2024.
format Preprint
id arxiv_https___arxiv_org_abs_2405_04485
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adapting WavLM for Speech Emotion Recognition
Diatlova, Daria
Udalov, Anton
Shutov, Vitalii
Spirin, Egor
Machine Learning
Sound
Audio and Speech Processing
Recently, the usage of speech self-supervised models (SSL) for downstream tasks has been drawing a lot of attention. While large pre-trained models commonly outperform smaller models trained from scratch, questions regarding the optimal fine-tuning strategies remain prevalent. In this paper, we explore the fine-tuning strategies of the WavLM Large model for the speech emotion recognition task on the MSP Podcast Corpus. More specifically, we perform a series of experiments focusing on using gender and semantic information from utterances. We then sum up our findings and describe the final model we used for submission to Speech Emotion Recognition Challenge 2024.
title Adapting WavLM for Speech Emotion Recognition
topic Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2405.04485