Scalable Controllable Accented TTS
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866918121125707776 |
|---|---|
| author | Xinyuan, Henry Li Cai, Zexin Garg, Ashi Duh, Kevin García-Perera, Leibny Paola Khudanpur, Sanjeev Andrews, Nicholas Wiesner, Matthew |
| author_facet | Xinyuan, Henry Li Cai, Zexin Garg, Ashi Duh, Kevin García-Perera, Leibny Paola Khudanpur, Sanjeev Andrews, Nicholas Wiesner, Matthew |
| contents | We tackle the challenge of scaling accented TTS systems, expanding their capabilities to include much larger amounts of training data and a wider variety of accent labels, even for accents that are poorly represented or unlabeled in traditional TTS datasets. To achieve this, we employ two strategies: 1. Accent label discovery via a speech geolocation model, which automatically infers accent labels from raw speech data without relying solely on human annotation; 2. Timbre augmentation through kNN voice conversion to increase data diversity and model robustness. These strategies are validated on CommonVoice, where we fine-tune XTTS-v2 for accented TTS with accent labels discovered or enhanced using geolocation. We demonstrate that the resulting accented TTS model not only outperforms XTTS-v2 fine-tuned on self-reported accent labels in CommonVoice, but also existing accented TTS benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_07426 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Scalable Controllable Accented TTS Xinyuan, Henry Li Cai, Zexin Garg, Ashi Duh, Kevin García-Perera, Leibny Paola Khudanpur, Sanjeev Andrews, Nicholas Wiesner, Matthew Audio and Speech Processing We tackle the challenge of scaling accented TTS systems, expanding their capabilities to include much larger amounts of training data and a wider variety of accent labels, even for accents that are poorly represented or unlabeled in traditional TTS datasets. To achieve this, we employ two strategies: 1. Accent label discovery via a speech geolocation model, which automatically infers accent labels from raw speech data without relying solely on human annotation; 2. Timbre augmentation through kNN voice conversion to increase data diversity and model robustness. These strategies are validated on CommonVoice, where we fine-tune XTTS-v2 for accented TTS with accent labels discovered or enhanced using geolocation. We demonstrate that the resulting accented TTS model not only outperforms XTTS-v2 fine-tuned on self-reported accent labels in CommonVoice, but also existing accented TTS benchmarks. |
| title | Scalable Controllable Accented TTS |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2508.07426 |