BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsu, Chan-Jan, Lin, Yi-Cheng, Lin, Chia-Chun, Chen, Wei-Chih, Chung, Ho Lam, Li, Chen-An, Chen, Yi-Chang, Yu, Chien-Yu, Lee, Ming-Ji, Chen, Chien-Cheng, Huang, Ru-Heng, Lee, Hung-yi, Shiu, Da-Shan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910804846051328
author Hsu, Chan-Jan
Lin, Yi-Cheng
Lin, Chia-Chun
Chen, Wei-Chih
Chung, Ho Lam
Li, Chen-An
Chen, Yi-Chang
Yu, Chien-Yu
Lee, Ming-Ji
Chen, Chien-Cheng
Huang, Ru-Heng
Lee, Hung-yi
Shiu, Da-Shan
author_facet Hsu, Chan-Jan
Lin, Yi-Cheng
Lin, Chia-Chun
Chen, Wei-Chih
Chung, Ho Lam
Li, Chen-An
Chen, Yi-Chang
Yu, Chien-Yu
Lee, Ming-Ji
Chen, Chien-Cheng
Huang, Ru-Heng
Lee, Hung-yi
Shiu, Da-Shan
contents We present BreezyVoice, a Text-to-Speech (TTS) system specifically adapted for Taiwanese Mandarin, highlighting phonetic control abilities to address the unique challenges of polyphone disambiguation in the language. Building upon CosyVoice, we incorporate a $S^{3}$ tokenizer, a large language model (LLM), an optimal-transport conditional flow matching model (OT-CFM), and a grapheme to phoneme prediction model, to generate realistic speech that closely mimics human utterances. Our evaluation demonstrates BreezyVoice's superior performance in both general and code-switching contexts, highlighting its robustness and effectiveness in generating high-fidelity speech. Additionally, we address the challenges of generalizability in modeling long-tail speakers and polyphone disambiguation. Our approach significantly enhances performance and offers valuable insights into the workings of neural codec TTS systems.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17790
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights
Hsu, Chan-Jan
Lin, Yi-Cheng
Lin, Chia-Chun
Chen, Wei-Chih
Chung, Ho Lam
Li, Chen-An
Chen, Yi-Chang
Yu, Chien-Yu
Lee, Ming-Ji
Chen, Chien-Cheng
Huang, Ru-Heng
Lee, Hung-yi
Shiu, Da-Shan
Computation and Language
Artificial Intelligence
We present BreezyVoice, a Text-to-Speech (TTS) system specifically adapted for Taiwanese Mandarin, highlighting phonetic control abilities to address the unique challenges of polyphone disambiguation in the language. Building upon CosyVoice, we incorporate a $S^{3}$ tokenizer, a large language model (LLM), an optimal-transport conditional flow matching model (OT-CFM), and a grapheme to phoneme prediction model, to generate realistic speech that closely mimics human utterances. Our evaluation demonstrates BreezyVoice's superior performance in both general and code-switching contexts, highlighting its robustness and effectiveness in generating high-fidelity speech. Additionally, we address the challenges of generalizability in modeling long-tail speakers and polyphone disambiguation. Our approach significantly enhances performance and offers valuable insights into the workings of neural codec TTS systems.
title BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.17790