Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bae, Jaesung, Zheng, Xiuwen, Kim, Minje, Yoo, Chang D., Hasegawa-Johnson, Mark
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914523077672960
author Bae, Jaesung
Zheng, Xiuwen
Kim, Minje
Yoo, Chang D.
Hasegawa-Johnson, Mark
author_facet Bae, Jaesung
Zheng, Xiuwen
Kim, Minje
Yoo, Chang D.
Hasegawa-Johnson, Mark
contents Dysarthric speech quality assessment (DSQA) is critical for clinical diagnostics and inclusive speech technologies. However, subjective evaluation is costly and difficult to scale, and the scarcity of labeled data limits robust objective modeling. To address this, we propose a three-stage framework that leverages unlabeled dysarthric speech and large-scale typical speech datasets to scale training. A teacher model first generates pseudo-labels for unlabeled samples, followed by weakly supervised pretraining using a label-aware contrastive learning strategy that exposes the model to diverse speakers and acoustic conditions. The pretrained model is then fine-tuned for the downstream DSQA task. Experiments on five unseen datasets spanning multiple etiologies and languages demonstrate the robustness of our approach. Our Whisper-based baseline significantly outperforms SOTA DSQA predictors such as SpICE, and the full framework achieves an average SRCC of 0.761 across unseen test datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15988
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech
Bae, Jaesung
Zheng, Xiuwen
Kim, Minje
Yoo, Chang D.
Hasegawa-Johnson, Mark
Audio and Speech Processing
Artificial Intelligence
Machine Learning
Dysarthric speech quality assessment (DSQA) is critical for clinical diagnostics and inclusive speech technologies. However, subjective evaluation is costly and difficult to scale, and the scarcity of labeled data limits robust objective modeling. To address this, we propose a three-stage framework that leverages unlabeled dysarthric speech and large-scale typical speech datasets to scale training. A teacher model first generates pseudo-labels for unlabeled samples, followed by weakly supervised pretraining using a label-aware contrastive learning strategy that exposes the model to diverse speakers and acoustic conditions. The pretrained model is then fine-tuned for the downstream DSQA task. Experiments on five unseen datasets spanning multiple etiologies and languages demonstrate the robustness of our approach. Our Whisper-based baseline significantly outperforms SOTA DSQA predictors such as SpICE, and the full framework achieves an average SRCC of 0.761 across unseen test datasets.
title Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech
topic Audio and Speech Processing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.15988