CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Jinyuan, Lan, Tian, Yu, Xintao, He, Xue, Zhang, Hezhi, Wang, Ying, Magistry, Pierre, Valette, Mathieu, Li, Lei
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916000514965504
author Xu, Jinyuan
Lan, Tian
Yu, Xintao
He, Xue
Zhang, Hezhi
Wang, Ying
Magistry, Pierre
Valette, Mathieu
Li, Lei
author_facet Xu, Jinyuan
Lan, Tian
Yu, Xintao
He, Xue
Zhang, Hezhi
Wang, Ying
Magistry, Pierre
Valette, Mathieu
Li, Lei
contents Depression is a pressing global public health issue, yet publicly available Chinese-language resources for depression risk detection remain scarce and largely focus on binary classification. To address this limitation, we release CNSocialDepress, a benchmark dataset for depression risk detection on Chinese social media. The dataset contains 44,178 posts from 233 users; psychological experts annotated 10,306 depression-related segments. CNSocialDepress provides binary risk labels along with structured, multidimensional psychological attributes, enabling interpretable and fine-grained analyses of depressive signals. Experimental results demonstrate the dataset's utility across a range of NLP tasks, including structured psychological profiling and fine-tuning large language models for depression detection. Comprehensive evaluations highlight the dataset's effectiveness and practical value for depression risk identification and psychological analysis, thereby providing insights for mental health applications tailored to Chinese-speaking populations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11233
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis
Xu, Jinyuan
Lan, Tian
Yu, Xintao
He, Xue
Zhang, Hezhi
Wang, Ying
Magistry, Pierre
Valette, Mathieu
Li, Lei
Computation and Language
Depression is a pressing global public health issue, yet publicly available Chinese-language resources for depression risk detection remain scarce and largely focus on binary classification. To address this limitation, we release CNSocialDepress, a benchmark dataset for depression risk detection on Chinese social media. The dataset contains 44,178 posts from 233 users; psychological experts annotated 10,306 depression-related segments. CNSocialDepress provides binary risk labels along with structured, multidimensional psychological attributes, enabling interpretable and fine-grained analyses of depressive signals. Experimental results demonstrate the dataset's utility across a range of NLP tasks, including structured psychological profiling and fine-tuning large language models for depression detection. Comprehensive evaluations highlight the dataset's effectiveness and practical value for depression risk identification and psychological analysis, thereby providing insights for mental health applications tailored to Chinese-speaking populations.
title CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis
topic Computation and Language
url https://arxiv.org/abs/2510.11233