GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Seokgi, Kim, Jungjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909623014916096
author Lee, Seokgi
Kim, Jungjun
author_facet Lee, Seokgi
Kim, Jungjun
contents We present the gradual style adaptor TTS (GSA-TTS) with a novel style encoder that gradually encodes speaking styles from an acoustic reference for zero-shot speech synthesis. GSA first captures the local style of each semantic sound unit. Then the local styles are combined by self-attention to obtain a global style condition. This semantic and hierarchical encoding strategy provides a robust and rich style representation for an acoustic model. We test GSA-TTS on unseen speakers and obtain promising results regarding naturalness, speaker similarity, and intelligibility. Additionally, we explore the potential of GSA in terms of interpretability and controllability, which stems from its hierarchical structure.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19384
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
Lee, Seokgi
Kim, Jungjun
Computation and Language
Sound
Audio and Speech Processing
We present the gradual style adaptor TTS (GSA-TTS) with a novel style encoder that gradually encodes speaking styles from an acoustic reference for zero-shot speech synthesis. GSA first captures the local style of each semantic sound unit. Then the local styles are combined by self-attention to obtain a global style condition. This semantic and hierarchical encoding strategy provides a robust and rich style representation for an acoustic model. We test GSA-TTS on unseen speakers and obtain promising results regarding naturalness, speaker similarity, and intelligibility. Additionally, we explore the potential of GSA in terms of interpretability and controllability, which stems from its hierarchical structure.
title GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.19384