Evaluating pretrained speech embedding systems for dysarthria detection across heterogenous datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wihlborg, Lovisa, Goodall, Jemima, Wheatley, David, Webber, Jacob J., Tam, Johnny, Weaver, Christine, Pal, Suvankar, Chandran, Siddharthan, Seth, Sohan, Watts, Oliver, Valentini-Botinhao, Cassia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917329127866368
author Wihlborg, Lovisa
Goodall, Jemima
Wheatley, David
Webber, Jacob J.
Tam, Johnny
Weaver, Christine
Pal, Suvankar
Chandran, Siddharthan
Seth, Sohan
Watts, Oliver
Valentini-Botinhao, Cassia
author_facet Wihlborg, Lovisa
Goodall, Jemima
Wheatley, David
Webber, Jacob J.
Tam, Johnny
Weaver, Christine
Pal, Suvankar
Chandran, Siddharthan
Seth, Sohan
Watts, Oliver
Valentini-Botinhao, Cassia
contents We present a comprehensive evaluation of pretrained speech embedding systems for the detection of dysarthric speech using existing accessible data. Dysarthric speech datasets are often small and can suffer from recording biases as well as data imbalance. To address these we selected a range of datasets covering related conditions and adopt the use of several cross-validations runs to estimate the chance level. To certify that results are above chance, we compare the distribution of scores across these runs against the distribution of scores of a carefully crafted null hypothesis. In this manner, we evaluate 17 publicly available speech embedding systems across 6 different datasets, reporting the cross-validation performance on each. We also report cross-dataset results derived when training with one particular dataset and testing with another. We observed that within-dataset results vary considerably depending on the dataset, regardless of the embedding used, raising questions about which datasets should be used for benchmarking. We found that cross-dataset accuracy is, as expected, lower than within-dataset, highlighting challenges in the generalization of the systems. These findings have important implications for the clinical validity of systems trained and tested on the same dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19946
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating pretrained speech embedding systems for dysarthria detection across heterogenous datasets
Wihlborg, Lovisa
Goodall, Jemima
Wheatley, David
Webber, Jacob J.
Tam, Johnny
Weaver, Christine
Pal, Suvankar
Chandran, Siddharthan
Seth, Sohan
Watts, Oliver
Valentini-Botinhao, Cassia
Audio and Speech Processing
We present a comprehensive evaluation of pretrained speech embedding systems for the detection of dysarthric speech using existing accessible data. Dysarthric speech datasets are often small and can suffer from recording biases as well as data imbalance. To address these we selected a range of datasets covering related conditions and adopt the use of several cross-validations runs to estimate the chance level. To certify that results are above chance, we compare the distribution of scores across these runs against the distribution of scores of a carefully crafted null hypothesis. In this manner, we evaluate 17 publicly available speech embedding systems across 6 different datasets, reporting the cross-validation performance on each. We also report cross-dataset results derived when training with one particular dataset and testing with another. We observed that within-dataset results vary considerably depending on the dataset, regardless of the embedding used, raising questions about which datasets should be used for benchmarking. We found that cross-dataset accuracy is, as expected, lower than within-dataset, highlighting challenges in the generalization of the systems. These findings have important implications for the clinical validity of systems trained and tested on the same dataset.
title Evaluating pretrained speech embedding systems for dysarthria detection across heterogenous datasets
topic Audio and Speech Processing
url https://arxiv.org/abs/2509.19946