Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xie, Xurong, Ruzi, Rukiye, Liu, Xunying, Wang, Lan
Formato: Preprint
Publicado: 2022
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929385802563584
author Xie, Xurong
Ruzi, Rukiye
Liu, Xunying
Wang, Lan
author_facet Xie, Xurong
Ruzi, Rukiye
Liu, Xunying
Wang, Lan
contents Dysarthric speech recognition is a challenging task due to acoustic variability and limited amount of available data. Diverse conditions of dysarthric speakers account for the acoustic variability, which make the variability difficult to be modeled precisely. This paper presents a variational auto-encoder based variability encoder (VAEVE) to explicitly encode such variability for dysarthric speech. The VAEVE makes use of both phoneme information and low-dimensional latent variable to reconstruct the input acoustic features, thereby the latent variable is forced to encode the phoneme-independent variability. Stochastic gradient variational Bayes algorithm is applied to model the distribution for generating variability encodings, which are further used as auxiliary features for DNN acoustic modeling. Experiment results conducted on the UASpeech corpus show that the VAEVE based variability encodings have complementary effect to the learning hidden unit contributions (LHUC) speaker adaptation. The systems using variability encodings consistently outperform the comparable baseline systems without using them, and" obtain absolute word error rate (WER) reduction by up to 2.2% on dysarthric speech with "Very lowintelligibility level, and up to 2% on the "Mixed" type of dysarthric speech with diverse or uncertain conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2201_09422
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
Xie, Xurong
Ruzi, Rukiye
Liu, Xunying
Wang, Lan
Audio and Speech Processing
Sound
Dysarthric speech recognition is a challenging task due to acoustic variability and limited amount of available data. Diverse conditions of dysarthric speakers account for the acoustic variability, which make the variability difficult to be modeled precisely. This paper presents a variational auto-encoder based variability encoder (VAEVE) to explicitly encode such variability for dysarthric speech. The VAEVE makes use of both phoneme information and low-dimensional latent variable to reconstruct the input acoustic features, thereby the latent variable is forced to encode the phoneme-independent variability. Stochastic gradient variational Bayes algorithm is applied to model the distribution for generating variability encodings, which are further used as auxiliary features for DNN acoustic modeling. Experiment results conducted on the UASpeech corpus show that the VAEVE based variability encodings have complementary effect to the learning hidden unit contributions (LHUC) speaker adaptation. The systems using variability encodings consistently outperform the comparable baseline systems without using them, and" obtain absolute word error rate (WER) reduction by up to 2.2% on dysarthric speech with "Very lowintelligibility level, and up to 2% on the "Mixed" type of dysarthric speech with diverse or uncertain conditions.
title Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2201.09422