STTATTS: Unified Speech-To-Text And Text-To-Speech Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Toyin, Hawau Olamide, Li, Hao, Aldarmaki, Hanan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909362690195456
author Toyin, Hawau Olamide
Li, Hao
Aldarmaki, Hanan
author_facet Toyin, Hawau Olamide
Li, Hao
Aldarmaki, Hanan
contents Speech recognition and speech synthesis models are typically trained separately, each with its own set of learning objectives, training data, and model parameters, resulting in two distinct large networks. We propose a parameter-efficient approach to learning ASR and TTS jointly via a multi-task learning objective and shared parameters. Our evaluation demonstrates that the performance of our multi-task model is comparable to that of individually trained models while significantly saving computational and memory costs ($\sim$50\% reduction in the total number of parameters required for the two tasks combined). We experiment with English as a resource-rich language, and Arabic as a relatively low-resource language due to shortage of TTS data. Our models are trained with publicly available data, and both the training code and model checkpoints are openly available for further research.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18607
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle STTATTS: Unified Speech-To-Text And Text-To-Speech Model
Toyin, Hawau Olamide
Li, Hao
Aldarmaki, Hanan
Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
Speech recognition and speech synthesis models are typically trained separately, each with its own set of learning objectives, training data, and model parameters, resulting in two distinct large networks. We propose a parameter-efficient approach to learning ASR and TTS jointly via a multi-task learning objective and shared parameters. Our evaluation demonstrates that the performance of our multi-task model is comparable to that of individually trained models while significantly saving computational and memory costs ($\sim$50\% reduction in the total number of parameters required for the two tasks combined). We experiment with English as a resource-rich language, and Arabic as a relatively low-resource language due to shortage of TTS data. Our models are trained with publicly available data, and both the training code and model checkpoints are openly available for further research.
title STTATTS: Unified Speech-To-Text And Text-To-Speech Model
topic Computation and Language
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.18607