Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908384491470848 |
|---|---|
| author | Chang, Yi Ren, Zhao Zhao, Zhonghao Nguyen, Thanh Tam Qian, Kun Schultz, Tanja Schuller, Björn W. |
| author_facet | Chang, Yi Ren, Zhao Zhao, Zhonghao Nguyen, Thanh Tam Qian, Kun Schultz, Tanja Schuller, Björn W. |
| contents | Speech emotion recognition (SER) plays a crucial role in human-computer interaction. The emergence of edge devices in the Internet of Things (IoT) presents challenges in constructing intricate deep learning models due to constraints in memory and computational resources. Moreover, emotional speech data often contains private information, raising concerns about privacy leakage during the deployment of SER models. To address these challenges, we propose a data distillation framework to facilitate efficient development of SER models in IoT applications using a synthesised, smaller, and distilled dataset. Our experiments demonstrate that the distilled dataset can be effectively utilised to train SER models with fixed initialisation, achieving performances comparable to those developed using the original full emotional speech dataset. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_15119 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation Chang, Yi Ren, Zhao Zhao, Zhonghao Nguyen, Thanh Tam Qian, Kun Schultz, Tanja Schuller, Björn W. Sound Artificial Intelligence Audio and Speech Processing Speech emotion recognition (SER) plays a crucial role in human-computer interaction. The emergence of edge devices in the Internet of Things (IoT) presents challenges in constructing intricate deep learning models due to constraints in memory and computational resources. Moreover, emotional speech data often contains private information, raising concerns about privacy leakage during the deployment of SER models. To address these challenges, we propose a data distillation framework to facilitate efficient development of SER models in IoT applications using a synthesised, smaller, and distilled dataset. Our experiments demonstrate that the distilled dataset can be effectively utilised to train SER models with fixed initialisation, achieving performances comparable to those developed using the original full emotional speech dataset. |
| title | Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation |
| topic | Sound Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2406.15119 |