Scalable Controllable Accented TTS

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xinyuan, Henry Li, Cai, Zexin, Garg, Ashi, Duh, Kevin, García-Perera, Leibny Paola, Khudanpur, Sanjeev, Andrews, Nicholas, Wiesner, Matthew
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918121125707776
author Xinyuan, Henry Li
Cai, Zexin
Garg, Ashi
Duh, Kevin
García-Perera, Leibny Paola
Khudanpur, Sanjeev
Andrews, Nicholas
Wiesner, Matthew
author_facet Xinyuan, Henry Li
Cai, Zexin
Garg, Ashi
Duh, Kevin
García-Perera, Leibny Paola
Khudanpur, Sanjeev
Andrews, Nicholas
Wiesner, Matthew
contents We tackle the challenge of scaling accented TTS systems, expanding their capabilities to include much larger amounts of training data and a wider variety of accent labels, even for accents that are poorly represented or unlabeled in traditional TTS datasets. To achieve this, we employ two strategies: 1. Accent label discovery via a speech geolocation model, which automatically infers accent labels from raw speech data without relying solely on human annotation; 2. Timbre augmentation through kNN voice conversion to increase data diversity and model robustness. These strategies are validated on CommonVoice, where we fine-tune XTTS-v2 for accented TTS with accent labels discovered or enhanced using geolocation. We demonstrate that the resulting accented TTS model not only outperforms XTTS-v2 fine-tuned on self-reported accent labels in CommonVoice, but also existing accented TTS benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07426
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scalable Controllable Accented TTS
Xinyuan, Henry Li
Cai, Zexin
Garg, Ashi
Duh, Kevin
García-Perera, Leibny Paola
Khudanpur, Sanjeev
Andrews, Nicholas
Wiesner, Matthew
Audio and Speech Processing
We tackle the challenge of scaling accented TTS systems, expanding their capabilities to include much larger amounts of training data and a wider variety of accent labels, even for accents that are poorly represented or unlabeled in traditional TTS datasets. To achieve this, we employ two strategies: 1. Accent label discovery via a speech geolocation model, which automatically infers accent labels from raw speech data without relying solely on human annotation; 2. Timbre augmentation through kNN voice conversion to increase data diversity and model robustness. These strategies are validated on CommonVoice, where we fine-tune XTTS-v2 for accented TTS with accent labels discovered or enhanced using geolocation. We demonstrate that the resulting accented TTS model not only outperforms XTTS-v2 fine-tuned on self-reported accent labels in CommonVoice, but also existing accented TTS benchmarks.
title Scalable Controllable Accented TTS
topic Audio and Speech Processing
url https://arxiv.org/abs/2508.07426