Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ferreira, Alef Iury Siqueira, Gris, Lucas Rafael, Filho, Alexandre Ferro, Ólives, Lucas, Ribeiro, Daniel, Fernando, Luiz, Lustosa, Fernanda, Tanaka, Rodrigo, de Oliveira, Frederico Santos, Filho, Arlindo Galvão
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915320012210176
author Ferreira, Alef Iury Siqueira
Gris, Lucas Rafael
Filho, Alexandre Ferro
Ólives, Lucas
Ribeiro, Daniel
Fernando, Luiz
Lustosa, Fernanda
Tanaka, Rodrigo
de Oliveira, Frederico Santos
Filho, Arlindo Galvão
author_facet Ferreira, Alef Iury Siqueira
Gris, Lucas Rafael
Filho, Alexandre Ferro
Ólives, Lucas
Ribeiro, Daniel
Fernando, Luiz
Lustosa, Fernanda
Tanaka, Rodrigo
de Oliveira, Frederico Santos
Filho, Arlindo Galvão
contents Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the INTERSPEECH 2025 Speech Emotion Recognition in Naturalistic Conditions Challenge, focusing on categorical emotion recognition. Our method combines state-of-the-art audio models with text features enriched by prosodic and spectral cues. In particular, we investigate the effectiveness of Fundamental Frequency (F0) quantization and the use of a pretrained audio tagging model. We also employ an ensemble model to improve robustness. On the official test set, our system achieved a Macro F1-score of 39.79% (42.20% on validation). Our results underscore the potential of these methods, and analysis of fusion techniques confirmed the effectiveness of Graph Attention Networks. Our source code is publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02088
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
Ferreira, Alef Iury Siqueira
Gris, Lucas Rafael
Filho, Alexandre Ferro
Ólives, Lucas
Ribeiro, Daniel
Fernando, Luiz
Lustosa, Fernanda
Tanaka, Rodrigo
de Oliveira, Frederico Santos
Filho, Arlindo Galvão
Sound
Computation and Language
Machine Learning
Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the INTERSPEECH 2025 Speech Emotion Recognition in Naturalistic Conditions Challenge, focusing on categorical emotion recognition. Our method combines state-of-the-art audio models with text features enriched by prosodic and spectral cues. In particular, we investigate the effectiveness of Fundamental Frequency (F0) quantization and the use of a pretrained audio tagging model. We also employ an ensemble model to improve robustness. On the official test set, our system achieved a Macro F1-score of 39.79% (42.20% on validation). Our results underscore the potential of these methods, and analysis of fusion techniques confirmed the effectiveness of Graph Attention Networks. Our source code is publicly available.
title Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
topic Sound
Computation and Language
Machine Learning
url https://arxiv.org/abs/2506.02088