MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Martinez-Lucas, Luz, Mote, Pravin, Naini, Abinay Reddy, Abdelwahab, Mohammed, Busso, Carlos
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918405856034816
author Martinez-Lucas, Luz
Mote, Pravin
Naini, Abinay Reddy
Abdelwahab, Mohammed
Busso, Carlos
author_facet Martinez-Lucas, Luz
Mote, Pravin
Naini, Abinay Reddy
Abdelwahab, Mohammed
Busso, Carlos
contents Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions conveyed through speech. While early SER systems relied on limited datasets and traditional machine learning models, recent deep learning approaches demand largescale, naturalistic emotional corpora. To address this need, we introduce the MSP-Conversation corpus: a dataset of more than 70 hours of conversational audio with time-continuous emotional annotations and detailed speaker diarizations. The time-continuous annotations capture the dynamic and contextdependent nature of emotional expression. The annotations in the corpus include fine-grained temporal traces of valence, arousal, and dominance. The audio data is sourced from publicly available podcasts and overlaps with a subset of the isolated speaking turns in the MSP-Podcast corpus to facilitate direct comparisons between annotation methods (i.e., in-context versus out-of-context annotations). The paper outlines the development of the corpus, annotation methodology, analyses of the annotations, and baseline SER experiments, establishing the MSP-Conversation corpus as a valuable resource for advancing research in dynamic SER in naturalistic settings.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22536
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition
Martinez-Lucas, Luz
Mote, Pravin
Naini, Abinay Reddy
Abdelwahab, Mohammed
Busso, Carlos
Audio and Speech Processing
Sound
Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions conveyed through speech. While early SER systems relied on limited datasets and traditional machine learning models, recent deep learning approaches demand largescale, naturalistic emotional corpora. To address this need, we introduce the MSP-Conversation corpus: a dataset of more than 70 hours of conversational audio with time-continuous emotional annotations and detailed speaker diarizations. The time-continuous annotations capture the dynamic and contextdependent nature of emotional expression. The annotations in the corpus include fine-grained temporal traces of valence, arousal, and dominance. The audio data is sourced from publicly available podcasts and overlaps with a subset of the isolated speaking turns in the MSP-Podcast corpus to facilitate direct comparisons between annotation methods (i.e., in-context versus out-of-context annotations). The paper outlines the development of the corpus, annotation methodology, analyses of the annotations, and baseline SER experiments, establishing the MSP-Conversation corpus as a valuable resource for advancing research in dynamic SER in naturalistic settings.
title MSP-Conversation: A Corpus for Naturalistic, Time-Continuous Emotion Recognition
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2603.22536