Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chang, Yi, Ren, Zhao, Zhao, Zhonghao, Nguyen, Thanh Tam, Qian, Kun, Schultz, Tanja, Schuller, Björn W.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908384491470848
author Chang, Yi
Ren, Zhao
Zhao, Zhonghao
Nguyen, Thanh Tam
Qian, Kun
Schultz, Tanja
Schuller, Björn W.
author_facet Chang, Yi
Ren, Zhao
Zhao, Zhonghao
Nguyen, Thanh Tam
Qian, Kun
Schultz, Tanja
Schuller, Björn W.
contents Speech emotion recognition (SER) plays a crucial role in human-computer interaction. The emergence of edge devices in the Internet of Things (IoT) presents challenges in constructing intricate deep learning models due to constraints in memory and computational resources. Moreover, emotional speech data often contains private information, raising concerns about privacy leakage during the deployment of SER models. To address these challenges, we propose a data distillation framework to facilitate efficient development of SER models in IoT applications using a synthesised, smaller, and distilled dataset. Our experiments demonstrate that the distilled dataset can be effectively utilised to train SER models with fixed initialisation, achieving performances comparable to those developed using the original full emotional speech dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2406_15119
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
Chang, Yi
Ren, Zhao
Zhao, Zhonghao
Nguyen, Thanh Tam
Qian, Kun
Schultz, Tanja
Schuller, Björn W.
Sound
Artificial Intelligence
Audio and Speech Processing
Speech emotion recognition (SER) plays a crucial role in human-computer interaction. The emergence of edge devices in the Internet of Things (IoT) presents challenges in constructing intricate deep learning models due to constraints in memory and computational resources. Moreover, emotional speech data often contains private information, raising concerns about privacy leakage during the deployment of SER models. To address these challenges, we propose a data distillation framework to facilitate efficient development of SER models in IoT applications using a synthesised, smaller, and distilled dataset. Our experiments demonstrate that the distilled dataset can be effectively utilised to train SER models with fixed initialisation, achieving performances comparable to those developed using the original full emotional speech dataset.
title Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2406.15119