Charting 15 years of progress in deep learning for speech emotion recognition: A replication study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Triantafyllopoulos, Andreas, Batliner, Anton, Schuller, Björn W.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918113688158208
author Triantafyllopoulos, Andreas
Batliner, Anton
Schuller, Björn W.
author_facet Triantafyllopoulos, Andreas
Batliner, Anton
Schuller, Björn W.
contents Speech emotion recognition (SER) has long benefited from the adoption of deep learning methodologies. Deeper models -- with more layers and more trainable parameters -- are generally perceived as being `better' by the SER community. This raises the question -- \emph{how much better} are modern-era deep neural networks compared to their earlier iterations? Beyond that, the more important question of how to move forward remains as poignant as ever. SER is far from a solved problem; therefore, identifying the most prominent avenues of future research is of paramount importance. In the present contribution, we attempt a quantification of progress in the 15 years of research beginning with the introduction of the landmark 2009 INTERSPEECH Emotion Challenge. We conduct a large scale investigation of model architectures, spanning both audio-based models that rely on speech inputs and text-baed models that rely solely on transcriptions. Our results point towards diminishing returns and a plateau after the recent introduction of transformer architectures. Moreover, we demonstrate how perceptions of progress are conditioned on the particular selection of models that are compared. Our findings have important repercussions about the state-of-the-art in SER research and the paths forward
format Preprint
id arxiv_https___arxiv_org_abs_2508_02448
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
Triantafyllopoulos, Andreas
Batliner, Anton
Schuller, Björn W.
Sound
Audio and Speech Processing
Speech emotion recognition (SER) has long benefited from the adoption of deep learning methodologies. Deeper models -- with more layers and more trainable parameters -- are generally perceived as being `better' by the SER community. This raises the question -- \emph{how much better} are modern-era deep neural networks compared to their earlier iterations? Beyond that, the more important question of how to move forward remains as poignant as ever. SER is far from a solved problem; therefore, identifying the most prominent avenues of future research is of paramount importance. In the present contribution, we attempt a quantification of progress in the 15 years of research beginning with the introduction of the landmark 2009 INTERSPEECH Emotion Challenge. We conduct a large scale investigation of model architectures, spanning both audio-based models that rely on speech inputs and text-baed models that rely solely on transcriptions. Our results point towards diminishing returns and a plateau after the recent introduction of transformer architectures. Moreover, we demonstrate how perceptions of progress are conditioned on the particular selection of models that are compared. Our findings have important repercussions about the state-of-the-art in SER research and the paths forward
title Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2508.02448