SpeechQE: Estimating the Quality of Direct Speech Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, HyoJung, Duh, Kevin, Carpuat, Marine
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929564126543872
author Han, HyoJung
Duh, Kevin
Carpuat, Marine
author_facet Han, HyoJung
Duh, Kevin
Carpuat, Marine
contents Recent advances in automatic quality estimation for machine translation have exclusively focused on written language, leaving the speech modality underexplored. In this work, we formulate the task of quality estimation for speech translation (SpeechQE), construct a benchmark, and evaluate a family of systems based on cascaded and end-to-end architectures. In this process, we introduce a novel end-to-end system leveraging pre-trained text LLM. Results suggest that end-to-end approaches are better suited to estimating the quality of direct speech translation than using quality estimation systems designed for text in cascaded systems. More broadly, we argue that quality estimation of speech translation needs to be studied as a separate problem from that of text, and release our data and models to guide further research in this space.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21485
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SpeechQE: Estimating the Quality of Direct Speech Translation
Han, HyoJung
Duh, Kevin
Carpuat, Marine
Computation and Language
Recent advances in automatic quality estimation for machine translation have exclusively focused on written language, leaving the speech modality underexplored. In this work, we formulate the task of quality estimation for speech translation (SpeechQE), construct a benchmark, and evaluate a family of systems based on cascaded and end-to-end architectures. In this process, we introduce a novel end-to-end system leveraging pre-trained text LLM. Results suggest that end-to-end approaches are better suited to estimating the quality of direct speech translation than using quality estimation systems designed for text in cascaded systems. More broadly, we argue that quality estimation of speech translation needs to be studied as a separate problem from that of text, and release our data and models to guide further research in this space.
title SpeechQE: Estimating the Quality of Direct Speech Translation
topic Computation and Language
url https://arxiv.org/abs/2410.21485