Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhou, Giulio, Lam, Tsz Kin, Birch, Alexandra, Haddow, Barry
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911769185746944
author Zhou, Giulio
Lam, Tsz Kin
Birch, Alexandra
Haddow, Barry
author_facet Zhou, Giulio
Lam, Tsz Kin
Birch, Alexandra
Haddow, Barry
contents Speech-to-Text Translation (S2TT) has typically been addressed with cascade systems, where speech recognition systems generate a transcription that is subsequently passed to a translation model. While there has been a growing interest in developing direct speech translation systems to avoid propagating errors and losing non-verbal content, prior work in direct S2TT has struggled to conclusively establish the advantages of integrating the acoustic signal directly into the translation process. This work proposes using contrastive evaluation to quantitatively measure the ability of direct S2TT systems to disambiguate utterances where prosody plays a crucial role. Specifically, we evaluated Korean-English translation systems on a test set containing wh-phrases, for which prosodic features are necessary to produce translations with the correct intent, whether it's a statement, a yes/no question, a wh-question, and more. Our results clearly demonstrate the value of direct translation systems over cascade translation models, with a notable 12.9% improvement in overall accuracy in ambiguous cases, along with up to a 15.6% increase in F1 scores for one of the major intent categories. To the best of our knowledge, this work stands as the first to provide quantitative evidence that direct S2TT models can effectively leverage prosody. The code for our evaluation is openly accessible and freely available for review and utilisation.
format Preprint
id arxiv_https___arxiv_org_abs_2402_00632
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases
Zhou, Giulio
Lam, Tsz Kin
Birch, Alexandra
Haddow, Barry
Computation and Language
Speech-to-Text Translation (S2TT) has typically been addressed with cascade systems, where speech recognition systems generate a transcription that is subsequently passed to a translation model. While there has been a growing interest in developing direct speech translation systems to avoid propagating errors and losing non-verbal content, prior work in direct S2TT has struggled to conclusively establish the advantages of integrating the acoustic signal directly into the translation process. This work proposes using contrastive evaluation to quantitatively measure the ability of direct S2TT systems to disambiguate utterances where prosody plays a crucial role. Specifically, we evaluated Korean-English translation systems on a test set containing wh-phrases, for which prosodic features are necessary to produce translations with the correct intent, whether it's a statement, a yes/no question, a wh-question, and more. Our results clearly demonstrate the value of direct translation systems over cascade translation models, with a notable 12.9% improvement in overall accuracy in ambiguous cases, along with up to a 15.6% increase in F1 scores for one of the major intent categories. To the best of our knowledge, this work stands as the first to provide quantitative evidence that direct S2TT models can effectively leverage prosody. The code for our evaluation is openly accessible and freely available for review and utilisation.
title Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases
topic Computation and Language
url https://arxiv.org/abs/2402.00632