KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhaolin, Liu, Yining, Liu, Danni, Nguyen, Tuan Nam, Ugan, Enes Yavuz, Dinh, Tu Anh, Mullov, Carlos, Waibel, Alexander, Niehues, Jan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918309506580480
author Li, Zhaolin
Liu, Yining
Liu, Danni
Nguyen, Tuan Nam
Ugan, Enes Yavuz
Dinh, Tu Anh
Mullov, Carlos
Waibel, Alexander
Niehues, Jan
author_facet Li, Zhaolin
Liu, Yining
Liu, Danni
Nguyen, Tuan Nam
Ugan, Enes Yavuz
Dinh, Tu Anh
Mullov, Carlos
Waibel, Alexander
Niehues, Jan
contents This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translation (MT) models, and end-to-end (E2E) Speech Translation (ST) systems for three language pairs: Bemba, North Levantine Arabic, and Tunisian Arabic into English. Building upon pre-trained models, we fine-tune our systems with different strategies to utilize resources efficiently. This study further explores system enhancement with synthetic data and model regularization. Specifically, we investigate MT-augmented ST by generating translations from ASR data using MT models. For North Levantine, which lacks parallel ST training data, a system trained solely on synthetic data slightly surpasses the cascaded system trained on real data. We also explore augmentation using text-to-speech models by generating synthetic speech from MT data, demonstrating the benefits of synthetic data in improving both ASR and ST performance for Bemba. Additionally, we apply intra-distillation to enhance model performance. Our experiments show that this approach consistently improves results across ASR, MT, and ST tasks, as well as across different pre-trained models. Finally, we apply Minimum Bayes Risk decoding to combine the cascaded and end-to-end systems, achieving an improvement of approximately 1.5 BLEU points.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19679
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
Li, Zhaolin
Liu, Yining
Liu, Danni
Nguyen, Tuan Nam
Ugan, Enes Yavuz
Dinh, Tu Anh
Mullov, Carlos
Waibel, Alexander
Niehues, Jan
Computation and Language
Artificial Intelligence
This paper presents KIT's submissions to the IWSLT 2025 low-resource track. We develop both cascaded systems, consisting of Automatic Speech Recognition (ASR) and Machine Translation (MT) models, and end-to-end (E2E) Speech Translation (ST) systems for three language pairs: Bemba, North Levantine Arabic, and Tunisian Arabic into English. Building upon pre-trained models, we fine-tune our systems with different strategies to utilize resources efficiently. This study further explores system enhancement with synthetic data and model regularization. Specifically, we investigate MT-augmented ST by generating translations from ASR data using MT models. For North Levantine, which lacks parallel ST training data, a system trained solely on synthetic data slightly surpasses the cascaded system trained on real data. We also explore augmentation using text-to-speech models by generating synthetic speech from MT data, demonstrating the benefits of synthetic data in improving both ASR and ST performance for Bemba. Additionally, we apply intra-distillation to enhance model performance. Our experiments show that this approach consistently improves results across ASR, MT, and ST tasks, as well as across different pre-trained models. Finally, we apply Minimum Bayes Risk decoding to combine the cascaded and end-to-end systems, achieving an improvement of approximately 1.5 BLEU points.
title KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.19679