How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Casals-Salvador, Marc, Costa, Federico, Zevallos, Rodolfo, Hernando, Javier
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915866734493696
author Casals-Salvador, Marc
Costa, Federico
Zevallos, Rodolfo
Hernando, Javier
author_facet Casals-Salvador, Marc
Costa, Federico
Zevallos, Rodolfo
Hernando, Javier
contents Speech Emotion Recognition (SER) plays a key role in advancing human-computer interaction. Attention mechanisms have become the dominant approach for modeling emotional speech due to their ability to capture long-range dependencies and emphasize salient information. However, standard self-attention suffers from quadratic computational and memory complexity, limiting its scalability. In this work, we present a systematic benchmark of optimized attention mechanisms for SER, including RetNet, LightNet, GSA, FoX, and KDA. Experiments on both MSP-Podcast benchmark versions show that while standard self-attention achieves the strongest recognition performance across test sets, efficient attention variants dramatically improve scalability, reducing inference latency and memory usage by up to an order of magnitude. These results highlight a critical trade-off between accuracy and efficiency, providing practical insights for designing scalable SER systems.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15120
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
Casals-Salvador, Marc
Costa, Federico
Zevallos, Rodolfo
Hernando, Javier
Audio and Speech Processing
Speech Emotion Recognition (SER) plays a key role in advancing human-computer interaction. Attention mechanisms have become the dominant approach for modeling emotional speech due to their ability to capture long-range dependencies and emphasize salient information. However, standard self-attention suffers from quadratic computational and memory complexity, limiting its scalability. In this work, we present a systematic benchmark of optimized attention mechanisms for SER, including RetNet, LightNet, GSA, FoX, and KDA. Experiments on both MSP-Podcast benchmark versions show that while standard self-attention achieves the strongest recognition performance across test sets, efficient attention variants dramatically improve scalability, reducing inference latency and memory usage by up to an order of magnitude. These results highlight a critical trade-off between accuracy and efficiency, providing practical insights for designing scalable SER systems.
title How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
topic Audio and Speech Processing
url https://arxiv.org/abs/2603.15120