Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bijoy, Mehedi Hasan, Porjazovski, Dejan, Grósz, Tamás, Kurimo, Mikko
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909999421194240
author Bijoy, Mehedi Hasan
Porjazovski, Dejan
Grósz, Tamás
Kurimo, Mikko
author_facet Bijoy, Mehedi Hasan
Porjazovski, Dejan
Grósz, Tamás
Kurimo, Mikko
contents Speech Emotion Recognition (SER) is crucial for improving human-computer interaction. Despite strides in monolingual SER, extending them to build a multilingual system remains challenging. Our goal is to train a single model capable of multilingual SER by distilling knowledge from multiple teacher models. To address this, we introduce a novel language-aware multi-teacher knowledge distillation method to advance SER in English, Finnish, and French. It leverages Wav2Vec2.0 as the foundation of monolingual teacher models and then distills their knowledge into a single multilingual student model. The student model demonstrates state-of-the-art performance, with a weighted recall of 72.9 on the English dataset and an unweighted recall of 63.4 on the Finnish dataset, surpassing fine-tuning and knowledge distillation baselines. Our method excels in improving recall for sad and neutral emotions, although it still faces challenges in recognizing anger and happiness.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08717
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
Bijoy, Mehedi Hasan
Porjazovski, Dejan
Grósz, Tamás
Kurimo, Mikko
Computation and Language
Sound
Audio and Speech Processing
Speech Emotion Recognition (SER) is crucial for improving human-computer interaction. Despite strides in monolingual SER, extending them to build a multilingual system remains challenging. Our goal is to train a single model capable of multilingual SER by distilling knowledge from multiple teacher models. To address this, we introduce a novel language-aware multi-teacher knowledge distillation method to advance SER in English, Finnish, and French. It leverages Wav2Vec2.0 as the foundation of monolingual teacher models and then distills their knowledge into a single multilingual student model. The student model demonstrates state-of-the-art performance, with a weighted recall of 72.9 on the English dataset and an unweighted recall of 63.4 on the Finnish dataset, surpassing fine-tuning and knowledge distillation baselines. Our method excels in improving recall for sad and neutral emotions, although it still faces challenges in recognizing anger and happiness.
title Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.08717