MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Inoue, Sho, Wang, Shuai, Wang, Wanxing, Zhu, Pengcheng, Bi, Mengxiao, Li, Haizhou
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929669401477120
author Inoue, Sho
Wang, Shuai
Wang, Wanxing
Zhu, Pengcheng
Bi, Mengxiao
Li, Haizhou
author_facet Inoue, Sho
Wang, Shuai
Wang, Wanxing
Zhu, Pengcheng
Bi, Mengxiao
Li, Haizhou
contents In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same speaker, through text transliteration for training accent conversion systems. We begin by generating transliterated text with Large Language Models (LLMs), which is then fed into multilingual TTS models to synthesize accented English speech. As a reference system, we built a sequence-to-sequence model on the synthetic parallel corpus for accent conversion. We validated the proposed method for both native and non-native English speakers. Subjective and objective evaluations further validate our dataset's effectiveness in accent conversion studies.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09352
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
Inoue, Sho
Wang, Shuai
Wang, Wanxing
Zhu, Pengcheng
Bi, Mengxiao
Li, Haizhou
Sound
Audio and Speech Processing
In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same speaker, through text transliteration for training accent conversion systems. We begin by generating transliterated text with Large Language Models (LLMs), which is then fed into multilingual TTS models to synthesize accented English speech. As a reference system, we built a sequence-to-sequence model on the synthetic parallel corpus for accent conversion. We validated the proposed method for both native and non-native English speakers. Subjective and objective evaluations further validate our dataset's effectiveness in accent conversion studies.
title MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.09352