Deep Speech Synthesis from Multimodal Articulatory Representations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wu, Peter, Yu, Bohan, Scheck, Kevin, Black, Alan W, Krishnapriyan, Aditi S., Chen, Irene Y., Schultz, Tanja, Watanabe, Shinji, Anumanchipalli, Gopala K.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910751291080704
author Wu, Peter
Yu, Bohan
Scheck, Kevin
Black, Alan W
Krishnapriyan, Aditi S.
Chen, Irene Y.
Schultz, Tanja
Watanabe, Shinji
Anumanchipalli, Gopala K.
author_facet Wu, Peter
Yu, Bohan
Scheck, Kevin
Black, Alan W
Krishnapriyan, Aditi S.
Chen, Irene Y.
Schultz, Tanja
Watanabe, Shinji
Anumanchipalli, Gopala K.
contents The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a multimodal pre-training framework. On single-speaker speech synthesis tasks from real-time magnetic resonance imaging and surface electromyography inputs, the intelligibility of synthesized outputs improves noticeably. For example, compared to prior work, utilizing our proposed transfer learning methods improves the MRI-to-speech performance by 36% word error rate. In addition to these intelligibility results, our multimodal pre-trained models consistently outperform unimodal baselines on three objective and subjective synthesis quality metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13387
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deep Speech Synthesis from Multimodal Articulatory Representations
Wu, Peter
Yu, Bohan
Scheck, Kevin
Black, Alan W
Krishnapriyan, Aditi S.
Chen, Irene Y.
Schultz, Tanja
Watanabe, Shinji
Anumanchipalli, Gopala K.
Audio and Speech Processing
Sound
The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a multimodal pre-training framework. On single-speaker speech synthesis tasks from real-time magnetic resonance imaging and surface electromyography inputs, the intelligibility of synthesized outputs improves noticeably. For example, compared to prior work, utilizing our proposed transfer learning methods improves the MRI-to-speech performance by 36% word error rate. In addition to these intelligibility results, our multimodal pre-trained models consistently outperform unimodal baselines on three objective and subjective synthesis quality metrics.
title Deep Speech Synthesis from Multimodal Articulatory Representations
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2412.13387