EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Haoqin, Wang, Xuechen, Zhao, Jinghua, Zhao, Shiwan, Zhou, Jiaming, Wang, Hui, He, Jiabei, Kong, Aobo, Yang, Xi, Wang, Yequan, Lin, Yonghua, Qin, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912448630489088
author Sun, Haoqin
Wang, Xuechen
Zhao, Jinghua
Zhao, Shiwan
Zhou, Jiaming
Wang, Hui
He, Jiabei
Kong, Aobo
Yang, Xi
Wang, Yequan
Lin, Yonghua
Qin, Yong
author_facet Sun, Haoqin
Wang, Xuechen
Zhao, Jinghua
Zhao, Shiwan
Zhou, Jiaming
Wang, Hui
He, Jiabei
Kong, Aobo
Yang, Xi
Wang, Yequan
Lin, Yonghua
Qin, Yong
contents In recent years, emotion recognition plays a critical role in applications such as human-computer interaction, mental health monitoring, and sentiment analysis. While datasets for emotion analysis in languages such as English have proliferated, there remains a pressing need for high-quality, comprehensive datasets tailored to the unique linguistic, cultural, and multimodal characteristics of Chinese. In this work, we propose \textbf{EmotionTalk}, an interactive Chinese multimodal emotion dataset with rich annotations. This dataset provides multimodal information from 19 actors participating in dyadic conversational settings, incorporating acoustic, visual, and textual modalities. It includes 23.6 hours of speech (19,250 utterances), annotations for 7 utterance-level emotion categories (happy, surprise, sad, disgust, anger, fear, and neutral), 5-dimensional sentiment labels (negative, weakly negative, neutral, weakly positive, and positive) and 4-dimensional speech captions (speaker, speaking style, emotion and overall). The dataset is well-suited for research on unimodal and multimodal emotion recognition, missing modality challenges, and speech captioning tasks. To our knowledge, it represents the first high-quality and versatile Chinese dialogue multimodal emotion dataset, which is a valuable contribution to research on cross-cultural emotion analysis and recognition. Additionally, we conduct experiments on EmotionTalk to demonstrate the effectiveness and quality of the dataset. It will be open-source and freely available for all academic purposes. The dataset and codes will be made available at: https://github.com/NKU-HLT/EmotionTalk.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23018
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
Sun, Haoqin
Wang, Xuechen
Zhao, Jinghua
Zhao, Shiwan
Zhou, Jiaming
Wang, Hui
He, Jiabei
Kong, Aobo
Yang, Xi
Wang, Yequan
Lin, Yonghua
Qin, Yong
Multimedia
In recent years, emotion recognition plays a critical role in applications such as human-computer interaction, mental health monitoring, and sentiment analysis. While datasets for emotion analysis in languages such as English have proliferated, there remains a pressing need for high-quality, comprehensive datasets tailored to the unique linguistic, cultural, and multimodal characteristics of Chinese. In this work, we propose \textbf{EmotionTalk}, an interactive Chinese multimodal emotion dataset with rich annotations. This dataset provides multimodal information from 19 actors participating in dyadic conversational settings, incorporating acoustic, visual, and textual modalities. It includes 23.6 hours of speech (19,250 utterances), annotations for 7 utterance-level emotion categories (happy, surprise, sad, disgust, anger, fear, and neutral), 5-dimensional sentiment labels (negative, weakly negative, neutral, weakly positive, and positive) and 4-dimensional speech captions (speaker, speaking style, emotion and overall). The dataset is well-suited for research on unimodal and multimodal emotion recognition, missing modality challenges, and speech captioning tasks. To our knowledge, it represents the first high-quality and versatile Chinese dialogue multimodal emotion dataset, which is a valuable contribution to research on cross-cultural emotion analysis and recognition. Additionally, we conduct experiments on EmotionTalk to demonstrate the effectiveness and quality of the dataset. It will be open-source and freely available for all academic purposes. The dataset and codes will be made available at: https://github.com/NKU-HLT/EmotionTalk.
title EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
topic Multimedia
url https://arxiv.org/abs/2505.23018