CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hira, Medha, Goel, Arnav, Gupta, Anubha
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917696422019072
author Hira, Medha
Goel, Arnav
Gupta, Anubha
author_facet Hira, Medha
Goel, Arnav
Gupta, Anubha
contents This paper presents CrossVoice, a novel cascade-based Speech-to-Speech Translation (S2ST) system employing advanced ASR, MT, and TTS technologies with cross-lingual prosody preservation through transfer learning. We conducted comprehensive experiments comparing CrossVoice with direct-S2ST systems, showing improved BLEU scores on tasks such as Fisher Es-En, VoxPopuli Fr-En and prosody preservation on benchmark datasets CVSS-T and IndicTTS. With an average mean opinion score of 3.75 out of 4, speech synthesized by CrossVoice closely rivals human speech on the benchmark, highlighting the efficacy of cascade-based systems and transfer learning in multilingual S2ST with prosody transfer.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00021
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
Hira, Medha
Goel, Arnav
Gupta, Anubha
Computation and Language
Sound
Audio and Speech Processing
This paper presents CrossVoice, a novel cascade-based Speech-to-Speech Translation (S2ST) system employing advanced ASR, MT, and TTS technologies with cross-lingual prosody preservation through transfer learning. We conducted comprehensive experiments comparing CrossVoice with direct-S2ST systems, showing improved BLEU scores on tasks such as Fisher Es-En, VoxPopuli Fr-En and prosody preservation on benchmark datasets CVSS-T and IndicTTS. With an average mean opinion score of 3.75 out of 4, speech synthesized by CrossVoice closely rivals human speech on the benchmark, highlighting the efficacy of cascade-based systems and transfer learning in multilingual S2ST with prosody transfer.
title CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2406.00021