Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pendyala, Varsha, Morgado, Pedro, Sethares, William
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912458757636096
author Pendyala, Varsha
Morgado, Pedro
Sethares, William
author_facet Pendyala, Varsha
Morgado, Pedro
Sethares, William
contents Voice interfaces integral to the human-computer interaction systems can benefit from speech emotion recognition (SER) to customize responses based on user emotions. Since humans convey emotions through multi-modal audio-visual cues, developing SER systems using both the modalities is beneficial. However, collecting a vast amount of labeled data for their development is expensive. This paper proposes a knowledge distillation framework called LightweightSER (LiSER) that leverages unlabeled audio-visual data for SER, using large teacher models built on advanced speech and face representation models. LiSER transfers knowledge regarding speech emotions and facial expressions from the teacher models to lightweight student models. Experiments conducted on two benchmark datasets, RAVDESS and CREMA-D, demonstrate that LiSER can reduce the dependence on extensive labeled datasets for SER tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2507_00055
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation
Pendyala, Varsha
Morgado, Pedro
Sethares, William
Machine Learning
Human-Computer Interaction
Multimedia
Audio and Speech Processing
Image and Video Processing
Signal Processing
Voice interfaces integral to the human-computer interaction systems can benefit from speech emotion recognition (SER) to customize responses based on user emotions. Since humans convey emotions through multi-modal audio-visual cues, developing SER systems using both the modalities is beneficial. However, collecting a vast amount of labeled data for their development is expensive. This paper proposes a knowledge distillation framework called LightweightSER (LiSER) that leverages unlabeled audio-visual data for SER, using large teacher models built on advanced speech and face representation models. LiSER transfers knowledge regarding speech emotions and facial expressions from the teacher models to lightweight student models. Experiments conducted on two benchmark datasets, RAVDESS and CREMA-D, demonstrate that LiSER can reduce the dependence on extensive labeled datasets for SER tasks.
title Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation
topic Machine Learning
Human-Computer Interaction
Multimedia
Audio and Speech Processing
Image and Video Processing
Signal Processing
url https://arxiv.org/abs/2507.00055