Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Falai, Alessio, Zhang, Ziyao, Gangoly, Akos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914004103856128
author Falai, Alessio
Zhang, Ziyao
Gangoly, Akos
author_facet Falai, Alessio
Zhang, Ziyao
Gangoly, Akos
contents In this paper we investigate cross-lingual Text-To-Speech (TTS) synthesis through the lens of adapters, in the context of lightweight TTS systems. In particular, we compare the tasks of unseen speaker and language adaptation with the goal of synthesising a target voice in a target language, in which the target voice has no recordings therein. Results from objective evaluations demonstrate the effectiveness of adapters in learning language-specific and speaker-specific information, allowing pre-trained models to learn unseen speaker identities or languages, while avoiding catastrophic forgetting of the original model's speaker or language information. Additionally, to measure how native the generated voices are in terms of accent, we propose and validate an objective metric inspired by mispronunciation detection techniques in second-language (L2) learners. The paper also provides insights into the impact of adapter placement, configuration and the number of speakers used.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18006
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
Falai, Alessio
Zhang, Ziyao
Gangoly, Akos
Audio and Speech Processing
Computation and Language
Machine Learning
Sound
In this paper we investigate cross-lingual Text-To-Speech (TTS) synthesis through the lens of adapters, in the context of lightweight TTS systems. In particular, we compare the tasks of unseen speaker and language adaptation with the goal of synthesising a target voice in a target language, in which the target voice has no recordings therein. Results from objective evaluations demonstrate the effectiveness of adapters in learning language-specific and speaker-specific information, allowing pre-trained models to learn unseen speaker identities or languages, while avoiding catastrophic forgetting of the original model's speaker or language information. Additionally, to measure how native the generated voices are in terms of accent, we propose and validate an objective metric inspired by mispronunciation detection techniques in second-language (L2) learners. The paper also provides insights into the impact of adapter placement, configuration and the number of speakers used.
title Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
topic Audio and Speech Processing
Computation and Language
Machine Learning
Sound
url https://arxiv.org/abs/2508.18006