Articulatory Feature Prediction from Surface EMG during Speech Production

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Jihwan, Huang, Kevin, Avramidis, Kleanthis, Pistrosch, Simon, Gonzalez-Machorro, Monica, Lee, Yoonjeong, Schuller, Björn, Goldstein, Louis, Narayanan, Shrikanth
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916765333716992
author Lee, Jihwan
Huang, Kevin
Avramidis, Kleanthis
Pistrosch, Simon
Gonzalez-Machorro, Monica
Lee, Yoonjeong
Schuller, Björn
Goldstein, Louis
Narayanan, Shrikanth
author_facet Lee, Jihwan
Huang, Kevin
Avramidis, Kleanthis
Pistrosch, Simon
Gonzalez-Machorro, Monica
Lee, Yoonjeong
Schuller, Björn
Goldstein, Louis
Narayanan, Shrikanth
contents We present a model for predicting articulatory features from surface electromyography (EMG) signals during speech production. The proposed model integrates convolutional layers and a Transformer block, followed by separate predictors for articulatory features. Our approach achieves a high prediction correlation of approximately 0.9 for most articulatory features. Furthermore, we demonstrate that these predicted articulatory features can be decoded into intelligible speech waveforms. To our knowledge, this is the first method to decode speech waveforms from surface EMG via articulatory features, offering a novel approach to EMG-based speech synthesis. Additionally, we analyze the relationship between EMG electrode placement and articulatory feature predictability, providing knowledge-driven insights for optimizing EMG electrode configurations. The source code and decoded speech samples are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13814
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Articulatory Feature Prediction from Surface EMG during Speech Production
Lee, Jihwan
Huang, Kevin
Avramidis, Kleanthis
Pistrosch, Simon
Gonzalez-Machorro, Monica
Lee, Yoonjeong
Schuller, Björn
Goldstein, Louis
Narayanan, Shrikanth
Audio and Speech Processing
Artificial Intelligence
Sound
We present a model for predicting articulatory features from surface electromyography (EMG) signals during speech production. The proposed model integrates convolutional layers and a Transformer block, followed by separate predictors for articulatory features. Our approach achieves a high prediction correlation of approximately 0.9 for most articulatory features. Furthermore, we demonstrate that these predicted articulatory features can be decoded into intelligible speech waveforms. To our knowledge, this is the first method to decode speech waveforms from surface EMG via articulatory features, offering a novel approach to EMG-based speech synthesis. Additionally, we analyze the relationship between EMG electrode placement and articulatory feature predictability, providing knowledge-driven insights for optimizing EMG electrode configurations. The source code and decoded speech samples are publicly available.
title Articulatory Feature Prediction from Surface EMG during Speech Production
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2505.13814