Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Hankun, Du, Chenpeng, Guo, Yiwei, Wang, Shuai, Chen, Xie, Yu, Kai
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913552199057408
author Wang, Hankun
Du, Chenpeng
Guo, Yiwei
Wang, Shuai
Chen, Xie
Yu, Kai
author_facet Wang, Hankun
Du, Chenpeng
Guo, Yiwei
Wang, Shuai
Chen, Xie
Yu, Kai
contents Recent popular decoder-only text-to-speech models are known for their ability of generating natural-sounding speech. However, such models sometimes suffer from word skipping and repeating due to the lack of explicit monotonic alignment constraints. In this paper, we notice from the attention maps that some particular attention heads of the decoder-only model indicate the alignments between speech and text. We call the attention maps of those heads Alignment-Emerged Attention Maps (AEAMs). Based on this discovery, we propose a novel inference method without altering the training process, named Attention-Constrained Inference (ACI), to facilitate monotonic synthesis. It first identifies AEAMs using the Attention Sweeping algorithm and then applies constraining masks on AEAMs. Our experimental results on decoder-only TTS model VALL-E show that the WER of synthesized speech is reduced by up to 20.5% relatively with ACI while the naturalness and speaker similarity are comparable.
format Preprint
id arxiv_https___arxiv_org_abs_2404_19723
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
Wang, Hankun
Du, Chenpeng
Guo, Yiwei
Wang, Shuai
Chen, Xie
Yu, Kai
Audio and Speech Processing
Sound
Recent popular decoder-only text-to-speech models are known for their ability of generating natural-sounding speech. However, such models sometimes suffer from word skipping and repeating due to the lack of explicit monotonic alignment constraints. In this paper, we notice from the attention maps that some particular attention heads of the decoder-only model indicate the alignments between speech and text. We call the attention maps of those heads Alignment-Emerged Attention Maps (AEAMs). Based on this discovery, we propose a novel inference method without altering the training process, named Attention-Constrained Inference (ACI), to facilitate monotonic synthesis. It first identifies AEAMs using the Attention Sweeping algorithm and then applies constraining masks on AEAMs. Our experimental results on decoder-only TTS model VALL-E show that the WER of synthesized speech is reduced by up to 20.5% relatively with ACI while the naturalness and speaker similarity are comparable.
title Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2404.19723