Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lei, Shun, Zhou, Yixuan, Chen, Liyang, Luo, Dan, Wu, Zhiyong, Wu, Xixin, Kang, Shiyin, Jiang, Tao, Zhou, Yahui, Han, Yuxing, Meng, Helen
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909163717656576
author Lei, Shun
Zhou, Yixuan
Chen, Liyang
Luo, Dan
Wu, Zhiyong
Wu, Xixin
Kang, Shiyin
Jiang, Tao
Zhou, Yahui
Han, Yuxing
Meng, Helen
author_facet Lei, Shun
Zhou, Yixuan
Chen, Liyang
Luo, Dan
Wu, Zhiyong
Wu, Xixin
Kang, Shiyin
Jiang, Tao
Zhou, Yahui
Han, Yuxing
Meng, Helen
contents Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker's voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language model-based TTS models show zero-shot speaker adaptation capabilities with only a 3-second acoustic prompt of an unseen speaker. However, they are limited by the length of the acoustic prompt, which makes it difficult to clone personal speaking style. In this paper, we propose a novel zero-shot TTS model with the multi-scale acoustic prompts based on a neural codec language model VALL-E. A speaker-aware text encoder is proposed to learn the personal speaking style at the phoneme-level from the style prompt consisting of multiple sentences. Following that, a VALL-E based acoustic decoder is utilized to model the timbre from the timbre prompt at the frame-level and generate speech. The experimental results show that our proposed method outperforms baselines in terms of naturalness and speaker similarity, and can achieve better performance by scaling out to a longer style prompt.
format Preprint
id arxiv_https___arxiv_org_abs_2309_11977
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
Lei, Shun
Zhou, Yixuan
Chen, Liyang
Luo, Dan
Wu, Zhiyong
Wu, Xixin
Kang, Shiyin
Jiang, Tao
Zhou, Yahui
Han, Yuxing
Meng, Helen
Sound
Audio and Speech Processing
Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker's voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language model-based TTS models show zero-shot speaker adaptation capabilities with only a 3-second acoustic prompt of an unseen speaker. However, they are limited by the length of the acoustic prompt, which makes it difficult to clone personal speaking style. In this paper, we propose a novel zero-shot TTS model with the multi-scale acoustic prompts based on a neural codec language model VALL-E. A speaker-aware text encoder is proposed to learn the personal speaking style at the phoneme-level from the style prompt consisting of multiple sentences. Following that, a VALL-E based acoustic decoder is utilized to model the timbre from the timbre prompt at the frame-level and generate speech. The experimental results show that our proposed method outperforms baselines in terms of naturalness and speaker similarity, and can achieve better performance by scaling out to a longer style prompt.
title Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2309.11977