DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jinhyeok, Lee, Junhyeok, Choi, Hyeong-Seok, Ji, Seunghun, Kim, Hyeongju, Lee, Juheon
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916370560581632
author Yang, Jinhyeok
Lee, Junhyeok
Choi, Hyeong-Seok
Ji, Seunghun
Kim, Hyeongju
Lee, Juheon
author_facet Yang, Jinhyeok
Lee, Junhyeok
Choi, Hyeong-Seok
Ji, Seunghun
Kim, Hyeongju
Lee, Juheon
contents Text-to-Speech (TTS) models have advanced significantly, aiming to accurately replicate human speech's diversity, including unique speaker identities and linguistic nuances. Despite these advancements, achieving an optimal balance between speaker-fidelity and text-intelligibility remains a challenge, particularly when diverse control demands are considered. Addressing this, we introduce DualSpeech, a TTS model that integrates phoneme-level latent diffusion with dual classifier-free guidance. This approach enables exceptional control over speaker-fidelity and text-intelligibility. Experimental results demonstrate that by utilizing the sophisticated control, DualSpeech surpasses existing state-of-the-art TTS models in performance. Demos are available at https://bit.ly/48Ewoib.
format Preprint
id arxiv_https___arxiv_org_abs_2408_14423
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
Yang, Jinhyeok
Lee, Junhyeok
Choi, Hyeong-Seok
Ji, Seunghun
Kim, Hyeongju
Lee, Juheon
Audio and Speech Processing
Sound
Text-to-Speech (TTS) models have advanced significantly, aiming to accurately replicate human speech's diversity, including unique speaker identities and linguistic nuances. Despite these advancements, achieving an optimal balance between speaker-fidelity and text-intelligibility remains a challenge, particularly when diverse control demands are considered. Addressing this, we introduce DualSpeech, a TTS model that integrates phoneme-level latent diffusion with dual classifier-free guidance. This approach enables exceptional control over speaker-fidelity and text-intelligibility. Experimental results demonstrate that by utilizing the sophisticated control, DualSpeech surpasses existing state-of-the-art TTS models in performance. Demos are available at https://bit.ly/48Ewoib.
title DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2408.14423