Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909231639166976 |
|---|---|
| author | Neekhara, Paarth Hussain, Shehzeen Ghosh, Subhankar Li, Jason Valle, Rafael Badlani, Rohan Ginsburg, Boris |
| author_facet | Neekhara, Paarth Hussain, Shehzeen Ghosh, Subhankar Li, Jason Valle, Rafael Badlani, Rohan Ginsburg, Boris |
| contents | Large Language Model (LLM) based text-to-speech (TTS) systems have demonstrated remarkable capabilities in handling large speech datasets and generating natural speech for new speakers. However, LLM-based TTS models are not robust as the generated output can contain repeating words, missing words and mis-aligned speech (referred to as hallucinations or attention errors), especially when the text contains multiple occurrences of the same token. We examine these challenges in an encoder-decoder transformer model and find that certain cross-attention heads in such models implicitly learn the text and speech alignment when trained for predicting speech tokens for a given text. To make the alignment more robust, we propose techniques utilizing CTC loss and attention priors that encourage monotonic cross-attention over the text tokens. Our guided attention training technique does not introduce any new learnable parameters and significantly improves robustness of LLM-based TTS models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_17957 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment Neekhara, Paarth Hussain, Shehzeen Ghosh, Subhankar Li, Jason Valle, Rafael Badlani, Rohan Ginsburg, Boris Sound Artificial Intelligence Audio and Speech Processing Large Language Model (LLM) based text-to-speech (TTS) systems have demonstrated remarkable capabilities in handling large speech datasets and generating natural speech for new speakers. However, LLM-based TTS models are not robust as the generated output can contain repeating words, missing words and mis-aligned speech (referred to as hallucinations or attention errors), especially when the text contains multiple occurrences of the same token. We examine these challenges in an encoder-decoder transformer model and find that certain cross-attention heads in such models implicitly learn the text and speech alignment when trained for predicting speech tokens for a given text. To make the alignment more robust, we propose techniques utilizing CTC loss and attention priors that encourage monotonic cross-attention over the text tokens. Our guided attention training technique does not introduce any new learnable parameters and significantly improves robustness of LLM-based TTS models. |
| title | Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment |
| topic | Sound Artificial Intelligence Audio and Speech Processing |
| url | https://arxiv.org/abs/2406.17957 |