Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Corrêa, Pedro, Lima, João, Moreno, Victor, Ueda, Lucas, Costa, Paula Dornhofer Paro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908619684970496
author Corrêa, Pedro
Lima, João
Moreno, Victor
Ueda, Lucas
Costa, Paula Dornhofer Paro
author_facet Corrêa, Pedro
Lima, João
Moreno, Victor
Ueda, Lucas
Costa, Paula Dornhofer Paro
contents Advancements in spoken language processing have driven the development of spoken language models (SLMs), designed to achieve universal audio understanding by jointly learning text and audio representations for a wide range of tasks. Although promising results have been achieved, there is growing discussion regarding these models' generalization capabilities and the extent to which they truly integrate audio and text modalities in their internal representations. In this work, we evaluate four SLMs on the task of speech emotion recognition using a dataset of emotionally incongruent speech samples, a condition under which the semantic content of the spoken utterance conveys one emotion while speech expressiveness conveys another. Our results indicate that SLMs rely predominantly on textual semantics rather than speech emotion to perform the task, indicating that text-related representations largely dominate over acoustic representations. We release both the code and the Emotionally Incongruent Synthetic Speech dataset (EMIS) to the community.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25054
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
Corrêa, Pedro
Lima, João
Moreno, Victor
Ueda, Lucas
Costa, Paula Dornhofer Paro
Computation and Language
Audio and Speech Processing
Advancements in spoken language processing have driven the development of spoken language models (SLMs), designed to achieve universal audio understanding by jointly learning text and audio representations for a wide range of tasks. Although promising results have been achieved, there is growing discussion regarding these models' generalization capabilities and the extent to which they truly integrate audio and text modalities in their internal representations. In this work, we evaluate four SLMs on the task of speech emotion recognition using a dataset of emotionally incongruent speech samples, a condition under which the semantic content of the spoken utterance conveys one emotion while speech expressiveness conveys another. Our results indicate that SLMs rely predominantly on textual semantics rather than speech emotion to perform the task, indicating that text-related representations largely dominate over acoustic representations. We release both the code and the Emotionally Incongruent Synthetic Speech dataset (EMIS) to the community.
title Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2510.25054