DepFlow: Disentangled Speech Generation to Mitigate Semantic Bias in Depression Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yuxin, Zhang, Xiangyu, Li, Yifei, Guo, Zhiwei, Zhang, Haoyang, Chng, Eng Siong, Guan, Cuntai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909979808628736
author Li, Yuxin
Zhang, Xiangyu
Li, Yifei
Guo, Zhiwei
Zhang, Haoyang
Chng, Eng Siong
Guan, Cuntai
author_facet Li, Yuxin
Zhang, Xiangyu
Li, Yifei
Guo, Zhiwei
Zhang, Haoyang
Chng, Eng Siong
Guan, Cuntai
contents Speech is a scalable and non-invasive biomarker for early mental health screening. However, widely used depression datasets like DAIC-WOZ exhibit strong coupling between linguistic sentiment and diagnostic labels, encouraging models to learn semantic shortcuts. As a result, model robustness may be compromised in real-world scenarios, such as Camouflaged Depression, where individuals maintain socially positive or neutral language despite underlying depressive states. To mitigate this semantic bias, we propose DepFlow, a three-stage depression-conditioned text-to-speech framework. First, a Depression Acoustic Encoder learns speaker- and content-invariant depression embeddings through adversarial training, achieving effective disentanglement while preserving depression discriminability (ROC-AUC: 0.693). Second, a flow-matching TTS model with FiLM modulation injects these embeddings into synthesis, enabling control over depressive severity while preserving content and speaker identity. Third, a prototype-based severity mapping mechanism provides smooth and interpretable manipulation across the depression continuum. Using DepFlow, we construct a Camouflage Depression-oriented Augmentation (CDoA) dataset that pairs depressed acoustic patterns with positive/neutral content from a sentiment-stratified text bank, creating acoustic-semantic mismatches underrepresented in natural data. Evaluated across three depression detection architectures, CDoA improves macro-F1 by 9%, 12%, and 5%, respectively, consistently outperforming conventional augmentation strategies in depression Detection. Beyond enhancing robustness, DepFlow provides a controllable synthesis platform for conversational systems and simulation-based evaluation, where real clinical data remains limited by ethical and coverage constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2601_00303
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DepFlow: Disentangled Speech Generation to Mitigate Semantic Bias in Depression Detection
Li, Yuxin
Zhang, Xiangyu
Li, Yifei
Guo, Zhiwei
Zhang, Haoyang
Chng, Eng Siong
Guan, Cuntai
Computation and Language
Artificial Intelligence
Speech is a scalable and non-invasive biomarker for early mental health screening. However, widely used depression datasets like DAIC-WOZ exhibit strong coupling between linguistic sentiment and diagnostic labels, encouraging models to learn semantic shortcuts. As a result, model robustness may be compromised in real-world scenarios, such as Camouflaged Depression, where individuals maintain socially positive or neutral language despite underlying depressive states. To mitigate this semantic bias, we propose DepFlow, a three-stage depression-conditioned text-to-speech framework. First, a Depression Acoustic Encoder learns speaker- and content-invariant depression embeddings through adversarial training, achieving effective disentanglement while preserving depression discriminability (ROC-AUC: 0.693). Second, a flow-matching TTS model with FiLM modulation injects these embeddings into synthesis, enabling control over depressive severity while preserving content and speaker identity. Third, a prototype-based severity mapping mechanism provides smooth and interpretable manipulation across the depression continuum. Using DepFlow, we construct a Camouflage Depression-oriented Augmentation (CDoA) dataset that pairs depressed acoustic patterns with positive/neutral content from a sentiment-stratified text bank, creating acoustic-semantic mismatches underrepresented in natural data. Evaluated across three depression detection architectures, CDoA improves macro-F1 by 9%, 12%, and 5%, respectively, consistently outperforming conventional augmentation strategies in depression Detection. Beyond enhancing robustness, DepFlow provides a controllable synthesis platform for conversational systems and simulation-based evaluation, where real clinical data remains limited by ethical and coverage constraints.
title DepFlow: Disentangled Speech Generation to Mitigate Semantic Bias in Depression Detection
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.00303