MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866929669401477120 |
|---|---|
| author | Inoue, Sho Wang, Shuai Wang, Wanxing Zhu, Pengcheng Bi, Mengxiao Li, Haizhou |
| author_facet | Inoue, Sho Wang, Shuai Wang, Wanxing Zhu, Pengcheng Bi, Mengxiao Li, Haizhou |
| contents | In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same speaker, through text transliteration for training accent conversion systems. We begin by generating transliterated text with Large Language Models (LLMs), which is then fed into multilingual TTS models to synthesize accented English speech. As a reference system, we built a sequence-to-sequence model on the synthetic parallel corpus for accent conversion. We validated the proposed method for both native and non-native English speakers. Subjective and objective evaluations further validate our dataset's effectiveness in accent conversion studies. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_09352 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion Inoue, Sho Wang, Shuai Wang, Wanxing Zhu, Pengcheng Bi, Mengxiao Li, Haizhou Sound Audio and Speech Processing In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same speaker, through text transliteration for training accent conversion systems. We begin by generating transliterated text with Large Language Models (LLMs), which is then fed into multilingual TTS models to synthesize accented English speech. As a reference system, we built a sequence-to-sequence model on the synthetic parallel corpus for accent conversion. We validated the proposed method for both native and non-native English speakers. Subjective and objective evaluations further validate our dataset's effectiveness in accent conversion studies. |
| title | MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion |
| topic | Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.09352 |