Deep Speech Synthesis from Multimodal Articulatory Representations
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866910751291080704 |
|---|---|
| author | Wu, Peter Yu, Bohan Scheck, Kevin Black, Alan W Krishnapriyan, Aditi S. Chen, Irene Y. Schultz, Tanja Watanabe, Shinji Anumanchipalli, Gopala K. |
| author_facet | Wu, Peter Yu, Bohan Scheck, Kevin Black, Alan W Krishnapriyan, Aditi S. Chen, Irene Y. Schultz, Tanja Watanabe, Shinji Anumanchipalli, Gopala K. |
| contents | The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a multimodal pre-training framework. On single-speaker speech synthesis tasks from real-time magnetic resonance imaging and surface electromyography inputs, the intelligibility of synthesized outputs improves noticeably. For example, compared to prior work, utilizing our proposed transfer learning methods improves the MRI-to-speech performance by 36% word error rate. In addition to these intelligibility results, our multimodal pre-trained models consistently outperform unimodal baselines on three objective and subjective synthesis quality metrics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_13387 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Deep Speech Synthesis from Multimodal Articulatory Representations Wu, Peter Yu, Bohan Scheck, Kevin Black, Alan W Krishnapriyan, Aditi S. Chen, Irene Y. Schultz, Tanja Watanabe, Shinji Anumanchipalli, Gopala K. Audio and Speech Processing Sound The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a multimodal pre-training framework. On single-speaker speech synthesis tasks from real-time magnetic resonance imaging and surface electromyography inputs, the intelligibility of synthesized outputs improves noticeably. For example, compared to prior work, utilizing our proposed transfer learning methods improves the MRI-to-speech performance by 36% word error rate. In addition to these intelligibility results, our multimodal pre-trained models consistently outperform unimodal baselines on three objective and subjective synthesis quality metrics. |
| title | Deep Speech Synthesis from Multimodal Articulatory Representations |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2412.13387 |