Domain Adapting Deep Reinforcement Learning for Real-world Speech Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rajapakshe, Thejan, Rana, Rajib, Khalifa, Sara, Schuller, Bjorn W.
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909440477757440
author Rajapakshe, Thejan
Rana, Rajib
Khalifa, Sara
Schuller, Bjorn W.
author_facet Rajapakshe, Thejan
Rana, Rajib
Khalifa, Sara
Schuller, Bjorn W.
contents Computers can understand and then engage with people in an emotionally intelligent way thanks to speech-emotion recognition (SER). However, the performance of SER in cross-corpus and real-world live data feed scenarios can be significantly improved. The inability to adapt an existing model to a new domain is one of the shortcomings of SER methods. To address this challenge, researchers have developed domain adaptation techniques that transfer knowledge learnt by a model across the domain. Although existing domain adaptation techniques have improved performances across domains, they can be improved to adapt to a real-world live data feed situation where a model can self-tune while deployed. In this paper, we present a deep reinforcement learning-based strategy (RL-DA) for adapting a pre-trained model to a real-world live data feed setting while interacting with the environment and collecting continual feedback. RL-DA is evaluated on SER tasks, including cross-corpus and cross-language domain adaption schema. Evaluation results show that in a live data feed setting, RL-DA outperforms a baseline strategy by 11% and 14% in cross-corpus and cross-language scenarios, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2207_12248
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Domain Adapting Deep Reinforcement Learning for Real-world Speech Emotion Recognition
Rajapakshe, Thejan
Rana, Rajib
Khalifa, Sara
Schuller, Bjorn W.
Sound
Machine Learning
Audio and Speech Processing
Computers can understand and then engage with people in an emotionally intelligent way thanks to speech-emotion recognition (SER). However, the performance of SER in cross-corpus and real-world live data feed scenarios can be significantly improved. The inability to adapt an existing model to a new domain is one of the shortcomings of SER methods. To address this challenge, researchers have developed domain adaptation techniques that transfer knowledge learnt by a model across the domain. Although existing domain adaptation techniques have improved performances across domains, they can be improved to adapt to a real-world live data feed situation where a model can self-tune while deployed. In this paper, we present a deep reinforcement learning-based strategy (RL-DA) for adapting a pre-trained model to a real-world live data feed setting while interacting with the environment and collecting continual feedback. RL-DA is evaluated on SER tasks, including cross-corpus and cross-language domain adaption schema. Evaluation results show that in a live data feed setting, RL-DA outperforms a baseline strategy by 11% and 14% in cross-corpus and cross-language scenarios, respectively.
title Domain Adapting Deep Reinforcement Learning for Real-world Speech Emotion Recognition
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2207.12248