MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chatzichristodoulou, Georgios, Kosmopoulou, Despoina, Kritikos, Antonios, Poulopoulou, Anastasia, Georgiou, Efthymios, Katsamanis, Athanasios, Katsouros, Vassilis, Potamianos, Alexandros
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911137940897792
author Chatzichristodoulou, Georgios
Kosmopoulou, Despoina
Kritikos, Antonios
Poulopoulou, Anastasia
Georgiou, Efthymios
Katsamanis, Athanasios
Katsouros, Vassilis
Potamianos, Alexandros
author_facet Chatzichristodoulou, Georgios
Kosmopoulou, Despoina
Kritikos, Antonios
Poulopoulou, Anastasia
Georgiou, Efthymios
Katsamanis, Athanasios
Katsouros, Vassilis
Potamianos, Alexandros
contents SER is a challenging task due to the subjective nature of human emotions and their uneven representation under naturalistic conditions. We propose MEDUSA, a multimodal framework with a four-stage training pipeline, which effectively handles class imbalance and emotion ambiguity. The first two stages train an ensemble of classifiers that utilize DeepSER, a novel extension of a deep cross-modal transformer fusion mechanism from pretrained self-supervised acoustic and linguistic representations. Manifold MixUp is employed for further regularization. The last two stages optimize a trainable meta-classifier that combines the ensemble predictions. Our training approach incorporates human annotation scores as soft targets, coupled with balanced data sampling and multitask learning. MEDUSA ranked 1st in Task 1: Categorical Emotion Recognition in the Interspeech 2025: Speech Emotion Recognition in Naturalistic Conditions Challenge.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09556
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions
Chatzichristodoulou, Georgios
Kosmopoulou, Despoina
Kritikos, Antonios
Poulopoulou, Anastasia
Georgiou, Efthymios
Katsamanis, Athanasios
Katsouros, Vassilis
Potamianos, Alexandros
Computation and Language
SER is a challenging task due to the subjective nature of human emotions and their uneven representation under naturalistic conditions. We propose MEDUSA, a multimodal framework with a four-stage training pipeline, which effectively handles class imbalance and emotion ambiguity. The first two stages train an ensemble of classifiers that utilize DeepSER, a novel extension of a deep cross-modal transformer fusion mechanism from pretrained self-supervised acoustic and linguistic representations. Manifold MixUp is employed for further regularization. The last two stages optimize a trainable meta-classifier that combines the ensemble predictions. Our training approach incorporates human annotation scores as soft targets, coupled with balanced data sampling and multitask learning. MEDUSA ranked 1st in Task 1: Categorical Emotion Recognition in the Interspeech 2025: Speech Emotion Recognition in Naturalistic Conditions Challenge.
title MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions
topic Computation and Language
url https://arxiv.org/abs/2506.09556