VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yeom, Jiheum, Kim, Heeseung, Choi, Jooyoung, Lee, Che Hyun, Park, Nohil, Yoon, Sungroh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910758100533248
author Yeom, Jiheum
Kim, Heeseung
Choi, Jooyoung
Lee, Che Hyun
Park, Nohil
Yoon, Sungroh
author_facet Yeom, Jiheum
Kim, Heeseung
Choi, Jooyoung
Lee, Che Hyun
Park, Nohil
Yoon, Sungroh
contents When applying parameter-efficient finetuning via LoRA onto speaker adaptive text-to-speech models, adaptation performance may decline compared to full-finetuned counterparts, especially for out-of-domain speakers. Here, we propose VoiceGuider, a parameter-efficient speaker adaptive text-to-speech system reinforced with autoguidance to enhance the speaker adaptation performance, reducing the gap against full-finetuned models. We carefully explore various ways of strengthening autoguidance, ultimately finding the optimal strategy. VoiceGuider as a result shows robust adaptation performance especially on extreme out-of-domain speech data. We provide audible samples in our demo page.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15759
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
Yeom, Jiheum
Kim, Heeseung
Choi, Jooyoung
Lee, Che Hyun
Park, Nohil
Yoon, Sungroh
Sound
Audio and Speech Processing
When applying parameter-efficient finetuning via LoRA onto speaker adaptive text-to-speech models, adaptation performance may decline compared to full-finetuned counterparts, especially for out-of-domain speakers. Here, we propose VoiceGuider, a parameter-efficient speaker adaptive text-to-speech system reinforced with autoguidance to enhance the speaker adaptation performance, reducing the gap against full-finetuned models. We carefully explore various ways of strengthening autoguidance, ultimately finding the optimal strategy. VoiceGuider as a result shows robust adaptation performance especially on extreme out-of-domain speech data. We provide audible samples in our demo page.
title VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.15759