Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Yerin, Lee, Jeehyun, Koo, Myoung-Wan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913599111299072
author Choi, Yerin
Lee, Jeehyun
Koo, Myoung-Wan
author_facet Choi, Yerin
Lee, Jeehyun
Koo, Myoung-Wan
contents Due to the subjective nature of current clinical evaluation, the need for automatic severity evaluation in dysarthric speech has emerged. DNN models outperform ML models but lack user-friendly explainability. ML models offer explainable results at a feature level, but their performance is comparatively lower. Current ML models extract various features from raw waveforms to predict severity. However, existing methods do not encompass all dysarthric features used in clinical evaluation. To address this gap, we propose a feature extraction method that minimizes information loss. We introduce an ASR transcription as a novel feature extraction source. We finetune the ASR model for dysarthric speech, then use this model to transcribe dysarthric speech and extract word segment boundary information. It enables capturing finer pronunciation and broader prosodic features. These features demonstrated an improved severity prediction performance to existing features: balanced accuracy of 83.72%.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03784
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
Choi, Yerin
Lee, Jeehyun
Koo, Myoung-Wan
Sound
Artificial Intelligence
Audio and Speech Processing
Due to the subjective nature of current clinical evaluation, the need for automatic severity evaluation in dysarthric speech has emerged. DNN models outperform ML models but lack user-friendly explainability. ML models offer explainable results at a feature level, but their performance is comparatively lower. Current ML models extract various features from raw waveforms to predict severity. However, existing methods do not encompass all dysarthric features used in clinical evaluation. To address this gap, we propose a feature extraction method that minimizes information loss. We introduce an ASR transcription as a novel feature extraction source. We finetune the ASR model for dysarthric speech, then use this model to transcribe dysarthric speech and extract word segment boundary information. It enables capturing finer pronunciation and broader prosodic features. These features demonstrated an improved severity prediction performance to existing features: balanced accuracy of 83.72%.
title Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2412.03784