Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Huimeng, Xie, Xurong, Geng, Mengzhe, Hu, Shujie, Xu, Haoning, Chen, Youjun, Li, Zhaoqing, Deng, Jiajun, Liu, Xunying
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916556172165120
author Wang, Huimeng
Xie, Xurong
Geng, Mengzhe
Hu, Shujie
Xu, Haoning
Chen, Youjun
Li, Zhaoqing
Deng, Jiajun
Liu, Xunying
author_facet Wang, Huimeng
Xie, Xurong
Geng, Mengzhe
Hu, Shujie
Xu, Haoning
Chen, Youjun
Li, Zhaoqing
Deng, Jiajun
Liu, Xunying
contents Discrete tokens extracted provide efficient and domain adaptable speech features. Their application to disordered speech that exhibits articulation imprecision and large mismatch against normal voice remains unexplored. To improve their phonetic discrimination that is weakened during unsupervised K-means or vector quantization of continuous features, this paper proposes novel phone-purity guided (PPG) discrete tokens for dysarthric speech recognition. Phonetic label supervision is used to regularize maximum likelihood and reconstruction error costs used in standard K-means and VAE-VQ based discrete token extraction. Experiments conducted on the UASpeech corpus suggest that the proposed PPG discrete token features extracted from HuBERT consistently outperform hybrid TDNN and End-to-End (E2E) Conformer systems using non-PPG based K-means or VAE-VQ tokens across varying codebook sizes by statistically significant word error rate (WER) reductions up to 0.99\% and 1.77\% absolute (3.21\% and 4.82\% relative) respectively on the UASpeech test set of 16 dysarthric speakers. The lowest WER of 23.25\% was obtained by combining systems using different token features. Consistent improvements on the phone purity metric were also achieved. T-SNE visualization further demonstrates sharper decision boundaries were produced between K-means/VAE-VQ clusters after introducing phone-purity guidance.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
Wang, Huimeng
Xie, Xurong
Geng, Mengzhe
Hu, Shujie
Xu, Haoning
Chen, Youjun
Li, Zhaoqing
Deng, Jiajun
Liu, Xunying
Sound
Audio and Speech Processing
Discrete tokens extracted provide efficient and domain adaptable speech features. Their application to disordered speech that exhibits articulation imprecision and large mismatch against normal voice remains unexplored. To improve their phonetic discrimination that is weakened during unsupervised K-means or vector quantization of continuous features, this paper proposes novel phone-purity guided (PPG) discrete tokens for dysarthric speech recognition. Phonetic label supervision is used to regularize maximum likelihood and reconstruction error costs used in standard K-means and VAE-VQ based discrete token extraction. Experiments conducted on the UASpeech corpus suggest that the proposed PPG discrete token features extracted from HuBERT consistently outperform hybrid TDNN and End-to-End (E2E) Conformer systems using non-PPG based K-means or VAE-VQ tokens across varying codebook sizes by statistically significant word error rate (WER) reductions up to 0.99\% and 1.77\% absolute (3.21\% and 4.82\% relative) respectively on the UASpeech test set of 16 dysarthric speakers. The lowest WER of 23.25\% was obtained by combining systems using different token features. Consistent improvements on the phone purity metric were also achieved. T-SNE visualization further demonstrates sharper decision boundaries were produced between K-means/VAE-VQ clusters after introducing phone-purity guidance.
title Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2501.04379