Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maharana, Sarthak Kumar, Adidam, Krishna Kamal, Nandi, Shoumik, Srivastava, Ajitesh
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917586690637824
author Maharana, Sarthak Kumar
Adidam, Krishna Kamal
Nandi, Shoumik
Srivastava, Ajitesh
author_facet Maharana, Sarthak Kumar
Adidam, Krishna Kamal
Nandi, Shoumik
Srivastava, Ajitesh
contents Acoustic-to-articulatory inversion (AAI) involves mapping from the acoustic to the articulatory space. Signal-processing features like the MFCCs, have been widely used for the AAI task. For subjects with dysarthric speech, AAI is challenging because of an imprecise and indistinct pronunciation. In this work, we perform AAI for dysarthric speech using representations from pre-trained self-supervised learning (SSL) models. We demonstrate the impact of different pre-trained features on this challenging AAI task, at low-resource conditions. In addition, we also condition x-vectors to the extracted SSL features to train a BLSTM network. In the seen case, we experiment with three AAI training schemes (subject-specific, pooled, and fine-tuned). The results, consistent across training schemes, reveal that DeCoAR, in the fine-tuned scheme, achieves a relative improvement of the Pearson Correlation Coefficient (CC) by ~1.81% and ~4.56% for healthy controls and patients, respectively, over MFCCs. We observe similar average trends for different SSL features in the unseen case. Overall, SSL networks like wav2vec, APC, and DeCoAR, trained with feature reconstruction or future timestep prediction tasks, perform well in predicting dysarthric articulatory trajectories.
format Preprint
id arxiv_https___arxiv_org_abs_2309_01108
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
Maharana, Sarthak Kumar
Adidam, Krishna Kamal
Nandi, Shoumik
Srivastava, Ajitesh
Audio and Speech Processing
Machine Learning
Sound
Acoustic-to-articulatory inversion (AAI) involves mapping from the acoustic to the articulatory space. Signal-processing features like the MFCCs, have been widely used for the AAI task. For subjects with dysarthric speech, AAI is challenging because of an imprecise and indistinct pronunciation. In this work, we perform AAI for dysarthric speech using representations from pre-trained self-supervised learning (SSL) models. We demonstrate the impact of different pre-trained features on this challenging AAI task, at low-resource conditions. In addition, we also condition x-vectors to the extracted SSL features to train a BLSTM network. In the seen case, we experiment with three AAI training schemes (subject-specific, pooled, and fine-tuned). The results, consistent across training schemes, reveal that DeCoAR, in the fine-tuned scheme, achieves a relative improvement of the Pearson Correlation Coefficient (CC) by ~1.81% and ~4.56% for healthy controls and patients, respectively, over MFCCs. We observe similar average trends for different SSL features in the unseen case. Overall, SSL networks like wav2vec, APC, and DeCoAR, trained with feature reconstruction or future timestep prediction tasks, perform well in predicting dysarthric articulatory trajectories.
title Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2309.01108