Articulatory Feature Prediction from Surface EMG during Speech Production
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916765333716992 |
|---|---|
| author | Lee, Jihwan Huang, Kevin Avramidis, Kleanthis Pistrosch, Simon Gonzalez-Machorro, Monica Lee, Yoonjeong Schuller, Björn Goldstein, Louis Narayanan, Shrikanth |
| author_facet | Lee, Jihwan Huang, Kevin Avramidis, Kleanthis Pistrosch, Simon Gonzalez-Machorro, Monica Lee, Yoonjeong Schuller, Björn Goldstein, Louis Narayanan, Shrikanth |
| contents | We present a model for predicting articulatory features from surface electromyography (EMG) signals during speech production. The proposed model integrates convolutional layers and a Transformer block, followed by separate predictors for articulatory features. Our approach achieves a high prediction correlation of approximately 0.9 for most articulatory features. Furthermore, we demonstrate that these predicted articulatory features can be decoded into intelligible speech waveforms. To our knowledge, this is the first method to decode speech waveforms from surface EMG via articulatory features, offering a novel approach to EMG-based speech synthesis. Additionally, we analyze the relationship between EMG electrode placement and articulatory feature predictability, providing knowledge-driven insights for optimizing EMG electrode configurations. The source code and decoded speech samples are publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_13814 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Articulatory Feature Prediction from Surface EMG during Speech Production Lee, Jihwan Huang, Kevin Avramidis, Kleanthis Pistrosch, Simon Gonzalez-Machorro, Monica Lee, Yoonjeong Schuller, Björn Goldstein, Louis Narayanan, Shrikanth Audio and Speech Processing Artificial Intelligence Sound We present a model for predicting articulatory features from surface electromyography (EMG) signals during speech production. The proposed model integrates convolutional layers and a Transformer block, followed by separate predictors for articulatory features. Our approach achieves a high prediction correlation of approximately 0.9 for most articulatory features. Furthermore, we demonstrate that these predicted articulatory features can be decoded into intelligible speech waveforms. To our knowledge, this is the first method to decode speech waveforms from surface EMG via articulatory features, offering a novel approach to EMG-based speech synthesis. Additionally, we analyze the relationship between EMG electrode placement and articulatory feature predictability, providing knowledge-driven insights for optimizing EMG electrode configurations. The source code and decoded speech samples are publicly available. |
| title | Articulatory Feature Prediction from Surface EMG during Speech Production |
| topic | Audio and Speech Processing Artificial Intelligence Sound |
| url | https://arxiv.org/abs/2505.13814 |