Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916650151837696 |
|---|---|
| author | Battenberg, Eric Skerry-Ryan, RJ Stanton, Daisy Mariooryad, Soroosh Shannon, Matt Salazar, Julian Kao, David |
| author_facet | Battenberg, Eric Skerry-Ryan, RJ Stanton, Daisy Mariooryad, Soroosh Shannon, Matt Salazar, Julian Kao, David |
| contents | Autoregressive (AR) Transformer-based sequence models are known to have difficulty generalizing to sequences longer than those seen during training. When applied to text-to-speech (TTS), these models tend to drop or repeat words or produce erratic output, especially for longer utterances. In this paper, we introduce enhancements aimed at AR Transformer-based encoder-decoder TTS systems that address these robustness and length generalization issues. Our approach uses an alignment mechanism to provide cross-attention operations with relative location information. The associated alignment position is learned as a latent property of the model via backpropagation and requires no external alignment information during training. While the approach is tailored to the monotonic nature of TTS input-output alignment, it is still able to benefit from the flexible modeling power of interleaved multi-head self- and cross-attention operations. A system incorporating these improvements, which we call Very Attentive Tacotron, matches the naturalness and expressiveness of a baseline T5-based TTS system, while eliminating problems with repeated or dropped words and enabling generalization to any practical utterance length. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_22179 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech Battenberg, Eric Skerry-Ryan, RJ Stanton, Daisy Mariooryad, Soroosh Shannon, Matt Salazar, Julian Kao, David Computation and Language Machine Learning Sound Audio and Speech Processing Autoregressive (AR) Transformer-based sequence models are known to have difficulty generalizing to sequences longer than those seen during training. When applied to text-to-speech (TTS), these models tend to drop or repeat words or produce erratic output, especially for longer utterances. In this paper, we introduce enhancements aimed at AR Transformer-based encoder-decoder TTS systems that address these robustness and length generalization issues. Our approach uses an alignment mechanism to provide cross-attention operations with relative location information. The associated alignment position is learned as a latent property of the model via backpropagation and requires no external alignment information during training. While the approach is tailored to the monotonic nature of TTS input-output alignment, it is still able to benefit from the flexible modeling power of interleaved multi-head self- and cross-attention operations. A system incorporating these improvements, which we call Very Attentive Tacotron, matches the naturalness and expressiveness of a baseline T5-based TTS system, while eliminating problems with repeated or dropped words and enabling generalization to any practical utterance length. |
| title | Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech |
| topic | Computation and Language Machine Learning Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.22179 |