Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hussain, Shehzeen, Neekhara, Paarth, Yang, Xuesong, Casanova, Edresson, Ghosh, Subhankar, Fejgin, Roy, Langman, Ryan, Desta, Mikyas, Tavabi, Leili, Li, Jason
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909807019032576
author Hussain, Shehzeen
Neekhara, Paarth
Yang, Xuesong
Casanova, Edresson
Ghosh, Subhankar
Fejgin, Roy
Langman, Ryan
Desta, Mikyas
Tavabi, Leili
Li, Jason
author_facet Hussain, Shehzeen
Neekhara, Paarth
Yang, Xuesong
Casanova, Edresson
Ghosh, Subhankar
Fejgin, Roy
Langman, Ryan
Desta, Mikyas
Tavabi, Leili
Li, Jason
contents Developing high-quality text-to-speech (TTS) systems for low-resource languages is challenging due to the scarcity of paired text and speech data. In contrast, automatic speech recognition (ASR) models for such languages are often more accessible, owing to large-scale multilingual pre-training efforts. We propose a framework based on Group Relative Policy Optimization (GRPO) to adapt an autoregressive, multilingual TTS model to new languages. Our method first establishes a language-agnostic foundation for TTS synthesis by training a multilingual baseline with International Phonetic Alphabet (IPA) tokens. Next, we fine-tune this model on limited paired data of the new languages to capture the target language's prosodic features. Finally, we apply GRPO to optimize the model using only unpaired text and speaker prompts, guided by a multi-objective reward from pretrained ASR, speaker verification, and audio quality estimation models. Experiments demonstrate that this pipeline produces intelligible and speaker-consistent speech in low-resource languages, substantially outperforming fine-tuning alone. Furthermore, our GRPO-based framework also improves TTS performance in high-resource languages, surpassing offline alignment methods such as Direct Preference Optimization (DPO) yielding superior intelligibility, speaker similarity, and audio quality.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21718
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
Hussain, Shehzeen
Neekhara, Paarth
Yang, Xuesong
Casanova, Edresson
Ghosh, Subhankar
Fejgin, Roy
Langman, Ryan
Desta, Mikyas
Tavabi, Leili
Li, Jason
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Developing high-quality text-to-speech (TTS) systems for low-resource languages is challenging due to the scarcity of paired text and speech data. In contrast, automatic speech recognition (ASR) models for such languages are often more accessible, owing to large-scale multilingual pre-training efforts. We propose a framework based on Group Relative Policy Optimization (GRPO) to adapt an autoregressive, multilingual TTS model to new languages. Our method first establishes a language-agnostic foundation for TTS synthesis by training a multilingual baseline with International Phonetic Alphabet (IPA) tokens. Next, we fine-tune this model on limited paired data of the new languages to capture the target language's prosodic features. Finally, we apply GRPO to optimize the model using only unpaired text and speaker prompts, guided by a multi-objective reward from pretrained ASR, speaker verification, and audio quality estimation models. Experiments demonstrate that this pipeline produces intelligible and speaker-consistent speech in low-resource languages, substantially outperforming fine-tuning alone. Furthermore, our GRPO-based framework also improves TTS performance in high-resource languages, surpassing offline alignment methods such as Direct Preference Optimization (DPO) yielding superior intelligibility, speaker similarity, and audio quality.
title Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization
topic Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2509.21718