SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bukhari, Hazim, Deshmukh, Soham, Dhamyal, Hira, Raj, Bhiksha, Singh, Rita
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917729904099328
author Bukhari, Hazim
Deshmukh, Soham
Dhamyal, Hira
Raj, Bhiksha
Singh, Rita
author_facet Bukhari, Hazim
Deshmukh, Soham
Dhamyal, Hira
Raj, Bhiksha
Singh, Rita
contents Speech Emotion Recognition (SER) has been traditionally formulated as a classification task. However, emotions are generally a spectrum whose distribution varies from situation to situation leading to poor Out-of-Domain (OOD) performance. We take inspiration from statistical formulation of Automatic Speech Recognition (ASR) and formulate the SER task as generating the most likely sequence of text tokens to infer emotion. The formulation breaks SER into predicting acoustic model features weighted by language model prediction. As an instance of this approach, we present SELM, an audio-conditioned language model for SER that predicts different emotion views. We train SELM on curated speech emotion corpus and test it on three OOD datasets (RAVDESS, CREMAD, IEMOCAP) not used in training. SELM achieves significant improvements over the state-of-the-art baselines, with 17% and 7% relative accuracy gains for RAVDESS and CREMA-D, respectively. Moreover, SELM can further boost its performance by Few-Shot Learning using a few annotated examples. The results highlight the effectiveness of our SER formulation, especially to improve performance in OOD scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15300
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
Bukhari, Hazim
Deshmukh, Soham
Dhamyal, Hira
Raj, Bhiksha
Singh, Rita
Sound
Audio and Speech Processing
Speech Emotion Recognition (SER) has been traditionally formulated as a classification task. However, emotions are generally a spectrum whose distribution varies from situation to situation leading to poor Out-of-Domain (OOD) performance. We take inspiration from statistical formulation of Automatic Speech Recognition (ASR) and formulate the SER task as generating the most likely sequence of text tokens to infer emotion. The formulation breaks SER into predicting acoustic model features weighted by language model prediction. As an instance of this approach, we present SELM, an audio-conditioned language model for SER that predicts different emotion views. We train SELM on curated speech emotion corpus and test it on three OOD datasets (RAVDESS, CREMAD, IEMOCAP) not used in training. SELM achieves significant improvements over the state-of-the-art baselines, with 17% and 7% relative accuracy gains for RAVDESS and CREMA-D, respectively. Moreover, SELM can further boost its performance by Few-Shot Learning using a few annotated examples. The results highlight the effectiveness of our SER formulation, especially to improve performance in OOD scenarios.
title SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2407.15300