Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Qingyu, Chen, Yushen, Niu, Zhikang, Wang, Chunhui, Yang, Yunting, Zhang, Bowen, Zhao, Jian, Zhu, Pengcheng, Yu, Kai, Chen, Xie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910015983452160
author Liu, Qingyu
Chen, Yushen
Niu, Zhikang
Wang, Chunhui
Yang, Yunting
Zhang, Bowen
Zhao, Jian
Zhu, Pengcheng
Yu, Kai
Chen, Xie
author_facet Liu, Qingyu
Chen, Yushen
Niu, Zhikang
Wang, Chunhui
Yang, Yunting
Zhang, Bowen
Zhao, Jian
Zhu, Pengcheng
Yu, Kai
Chen, Xie
contents Flow-matching-based text-to-speech (TTS) models have shown high-quality speech synthesis. However, most current flow-matching-based TTS models still rely on reference transcripts corresponding to the audio prompt for synthesis. This dependency prevents cross-lingual voice cloning when audio prompt transcripts are unavailable, particularly for unseen languages. The key challenges for flow-matching-based TTS models to remove audio prompt transcripts are identifying word boundaries during training and determining appropriate duration during inference. In this paper, we introduce Cross-Lingual F5-TTS, a framework that enables cross-lingual voice cloning without audio prompt transcripts. Our method preprocesses audio prompts by forced alignment to obtain word boundaries, enabling direct synthesis from audio prompts while excluding transcripts during training. To address the duration modeling challenge, we train speaking rate predictors at different linguistic granularities to derive duration from speaker pace. Experiments show that our approach matches the performance of F5-TTS while enabling cross-lingual voice cloning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14579
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
Liu, Qingyu
Chen, Yushen
Niu, Zhikang
Wang, Chunhui
Yang, Yunting
Zhang, Bowen
Zhao, Jian
Zhu, Pengcheng
Yu, Kai
Chen, Xie
Sound
Flow-matching-based text-to-speech (TTS) models have shown high-quality speech synthesis. However, most current flow-matching-based TTS models still rely on reference transcripts corresponding to the audio prompt for synthesis. This dependency prevents cross-lingual voice cloning when audio prompt transcripts are unavailable, particularly for unseen languages. The key challenges for flow-matching-based TTS models to remove audio prompt transcripts are identifying word boundaries during training and determining appropriate duration during inference. In this paper, we introduce Cross-Lingual F5-TTS, a framework that enables cross-lingual voice cloning without audio prompt transcripts. Our method preprocesses audio prompts by forced alignment to obtain word boundaries, enabling direct synthesis from audio prompts while excluding transcripts during training. To address the duration modeling challenge, we train speaking rate predictors at different linguistic granularities to derive duration from speaker pace. Experiments show that our approach matches the performance of F5-TTS while enabling cross-lingual voice cloning.
title Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
topic Sound
url https://arxiv.org/abs/2509.14579