Tracking Articulatory Dynamics in Speech with a Fixed-Weight BiLSTM-CNN Architecture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pillai, Leena G, Mubarak, D. Muhammad Noorul, Sherly, Elizabeth
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916707515236352
author Pillai, Leena G
Mubarak, D. Muhammad Noorul
Sherly, Elizabeth
author_facet Pillai, Leena G
Mubarak, D. Muhammad Noorul
Sherly, Elizabeth
contents Speech production is a complex sequential process which involve the coordination of various articulatory features. Among them tongue being a highly versatile active articulator responsible for shaping airflow to produce targeted speech sounds that are intellectual, clear, and distinct. This paper presents a novel approach for predicting tongue and lip articulatory features involved in a given speech acoustics using a stacked Bidirectional Long Short-Term Memory (BiLSTM) architecture, combined with a one-dimensional Convolutional Neural Network (CNN) for post-processing with fixed weights initialization. The proposed network is trained with two datasets consisting of simultaneously recorded speech and Electromagnetic Articulography (EMA) datasets, each introducing variations in terms of geographical origin, linguistic characteristics, phonetic diversity, and recording equipment. The performance of the model is assessed in Speaker Dependent (SD), Speaker Independent (SI), corpus dependent (CD) and cross corpus (CC) modes. Experimental results indicate that the proposed model with fixed weights approach outperformed the adaptive weights initialization with in relatively minimal number of training epochs. These findings contribute to the development of robust and efficient models for articulatory feature prediction, paving the way for advancements in speech production research and applications.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18099
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tracking Articulatory Dynamics in Speech with a Fixed-Weight BiLSTM-CNN Architecture
Pillai, Leena G
Mubarak, D. Muhammad Noorul
Sherly, Elizabeth
Sound
Computation and Language
Audio and Speech Processing
Speech production is a complex sequential process which involve the coordination of various articulatory features. Among them tongue being a highly versatile active articulator responsible for shaping airflow to produce targeted speech sounds that are intellectual, clear, and distinct. This paper presents a novel approach for predicting tongue and lip articulatory features involved in a given speech acoustics using a stacked Bidirectional Long Short-Term Memory (BiLSTM) architecture, combined with a one-dimensional Convolutional Neural Network (CNN) for post-processing with fixed weights initialization. The proposed network is trained with two datasets consisting of simultaneously recorded speech and Electromagnetic Articulography (EMA) datasets, each introducing variations in terms of geographical origin, linguistic characteristics, phonetic diversity, and recording equipment. The performance of the model is assessed in Speaker Dependent (SD), Speaker Independent (SI), corpus dependent (CD) and cross corpus (CC) modes. Experimental results indicate that the proposed model with fixed weights approach outperformed the adaptive weights initialization with in relatively minimal number of training epochs. These findings contribute to the development of robust and efficient models for articulatory feature prediction, paving the way for advancements in speech production research and applications.
title Tracking Articulatory Dynamics in Speech with a Fixed-Weight BiLSTM-CNN Architecture
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2504.18099