Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Neekhara, Paarth, Hussain, Shehzeen, Ghosh, Subhankar, Li, Jason, Valle, Rafael, Badlani, Rohan, Ginsburg, Boris
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909231639166976
author Neekhara, Paarth
Hussain, Shehzeen
Ghosh, Subhankar
Li, Jason
Valle, Rafael
Badlani, Rohan
Ginsburg, Boris
author_facet Neekhara, Paarth
Hussain, Shehzeen
Ghosh, Subhankar
Li, Jason
Valle, Rafael
Badlani, Rohan
Ginsburg, Boris
contents Large Language Model (LLM) based text-to-speech (TTS) systems have demonstrated remarkable capabilities in handling large speech datasets and generating natural speech for new speakers. However, LLM-based TTS models are not robust as the generated output can contain repeating words, missing words and mis-aligned speech (referred to as hallucinations or attention errors), especially when the text contains multiple occurrences of the same token. We examine these challenges in an encoder-decoder transformer model and find that certain cross-attention heads in such models implicitly learn the text and speech alignment when trained for predicting speech tokens for a given text. To make the alignment more robust, we propose techniques utilizing CTC loss and attention priors that encourage monotonic cross-attention over the text tokens. Our guided attention training technique does not introduce any new learnable parameters and significantly improves robustness of LLM-based TTS models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17957
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
Neekhara, Paarth
Hussain, Shehzeen
Ghosh, Subhankar
Li, Jason
Valle, Rafael
Badlani, Rohan
Ginsburg, Boris
Sound
Artificial Intelligence
Audio and Speech Processing
Large Language Model (LLM) based text-to-speech (TTS) systems have demonstrated remarkable capabilities in handling large speech datasets and generating natural speech for new speakers. However, LLM-based TTS models are not robust as the generated output can contain repeating words, missing words and mis-aligned speech (referred to as hallucinations or attention errors), especially when the text contains multiple occurrences of the same token. We examine these challenges in an encoder-decoder transformer model and find that certain cross-attention heads in such models implicitly learn the text and speech alignment when trained for predicting speech tokens for a given text. To make the alignment more robust, we propose techniques utilizing CTC loss and attention priors that encourage monotonic cross-attention over the text tokens. Our guided attention training technique does not introduce any new learnable parameters and significantly improves robustness of LLM-based TTS models.
title Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2406.17957