STTATTS: Unified Speech-To-Text And Text-To-Speech Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909362690195456 |
|---|---|
| author | Toyin, Hawau Olamide Li, Hao Aldarmaki, Hanan |
| author_facet | Toyin, Hawau Olamide Li, Hao Aldarmaki, Hanan |
| contents | Speech recognition and speech synthesis models are typically trained separately, each with its own set of learning objectives, training data, and model parameters, resulting in two distinct large networks. We propose a parameter-efficient approach to learning ASR and TTS jointly via a multi-task learning objective and shared parameters. Our evaluation demonstrates that the performance of our multi-task model is comparable to that of individually trained models while significantly saving computational and memory costs ($\sim$50\% reduction in the total number of parameters required for the two tasks combined). We experiment with English as a resource-rich language, and Arabic as a relatively low-resource language due to shortage of TTS data. Our models are trained with publicly available data, and both the training code and model checkpoints are openly available for further research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_18607 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | STTATTS: Unified Speech-To-Text And Text-To-Speech Model Toyin, Hawau Olamide Li, Hao Aldarmaki, Hanan Computation and Language Artificial Intelligence Machine Learning Sound Audio and Speech Processing Speech recognition and speech synthesis models are typically trained separately, each with its own set of learning objectives, training data, and model parameters, resulting in two distinct large networks. We propose a parameter-efficient approach to learning ASR and TTS jointly via a multi-task learning objective and shared parameters. Our evaluation demonstrates that the performance of our multi-task model is comparable to that of individually trained models while significantly saving computational and memory costs ($\sim$50\% reduction in the total number of parameters required for the two tasks combined). We experiment with English as a resource-rich language, and Arabic as a relatively low-resource language due to shortage of TTS data. Our models are trained with publicly available data, and both the training code and model checkpoints are openly available for further research. |
| title | STTATTS: Unified Speech-To-Text And Text-To-Speech Model |
| topic | Computation and Language Artificial Intelligence Machine Learning Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.18607 |