Finetuning End-to-End Models for Estonian Conversational Spoken Language Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sildam, Tiia, Velve, Andra, Alumäe, Tanel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916313262194688
author Sildam, Tiia
Velve, Andra
Alumäe, Tanel
author_facet Sildam, Tiia
Velve, Andra
Alumäe, Tanel
contents This paper investigates the finetuning of end-to-end models for bidirectional Estonian-English and Estonian-Russian conversational speech-to-text translation. Due to the limited availability of speech translation data for Estonian, we created additional training data by web scraping and synthesizing data from speech recognition datasets using machine translation. We evaluated three publicly available end-to-end models: Whisper, OWSM 3.1, and SeamlessM4T. Our results indicate that fine-tuning with synthetic data enhances translation accuracy by a large margin, with SeamlessM4T matching or surpassing cascaded speech translation systems that use state-of-the-art speech recognition and machine translation models.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03809
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Finetuning End-to-End Models for Estonian Conversational Spoken Language Translation
Sildam, Tiia
Velve, Andra
Alumäe, Tanel
Computation and Language
Audio and Speech Processing
This paper investigates the finetuning of end-to-end models for bidirectional Estonian-English and Estonian-Russian conversational speech-to-text translation. Due to the limited availability of speech translation data for Estonian, we created additional training data by web scraping and synthesizing data from speech recognition datasets using machine translation. We evaluated three publicly available end-to-end models: Whisper, OWSM 3.1, and SeamlessM4T. Our results indicate that fine-tuning with synthetic data enhances translation accuracy by a large margin, with SeamlessM4T matching or surpassing cascaded speech translation systems that use state-of-the-art speech recognition and machine translation models.
title Finetuning End-to-End Models for Estonian Conversational Spoken Language Translation
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2407.03809