High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Joun Yeop, Jeong, Myeonghun, Kim, Minchan, Lee, Ji-Hyun, Cho, Hoon-Young, Kim, Nam Soo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909230672379904
author Lee, Joun Yeop
Jeong, Myeonghun
Kim, Minchan
Lee, Ji-Hyun
Cho, Hoon-Young
Kim, Nam Soo
author_facet Lee, Joun Yeop
Jeong, Myeonghun
Kim, Minchan
Lee, Ji-Hyun
Cho, Hoon-Young
Kim, Nam Soo
contents We propose a novel two-stage text-to-speech (TTS) framework with two types of discrete tokens, i.e., semantic and acoustic tokens, for high-fidelity speech synthesis. It features two core components: the Interpreting module, which processes text and a speech prompt into semantic tokens focusing on linguistic contents and alignment, and the Speaking module, which captures the timbre of the target voice to generate acoustic tokens from semantic tokens, enriching speech reconstruction. The Interpreting stage employs a transducer for its robustness in aligning text to speech. In contrast, the Speaking stage utilizes a Conformer-based architecture integrated with a Grouped Masked Language Model (G-MLM) to boost computational efficiency. Our experiments verify that this innovative structure surpasses the conventional models in the zero-shot scenario in terms of speech quality and speaker similarity.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17310
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
Lee, Joun Yeop
Jeong, Myeonghun
Kim, Minchan
Lee, Ji-Hyun
Cho, Hoon-Young
Kim, Nam Soo
Audio and Speech Processing
We propose a novel two-stage text-to-speech (TTS) framework with two types of discrete tokens, i.e., semantic and acoustic tokens, for high-fidelity speech synthesis. It features two core components: the Interpreting module, which processes text and a speech prompt into semantic tokens focusing on linguistic contents and alignment, and the Speaking module, which captures the timbre of the target voice to generate acoustic tokens from semantic tokens, enriching speech reconstruction. The Interpreting stage employs a transducer for its robustness in aligning text to speech. In contrast, the Speaking stage utilizes a Conformer-based architecture integrated with a Grouped Masked Language Model (G-MLM) to boost computational efficiency. Our experiments verify that this innovative structure surpasses the conventional models in the zero-shot scenario in terms of speech quality and speaker similarity.
title High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
topic Audio and Speech Processing
url https://arxiv.org/abs/2406.17310