GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Chutong, Anastasopoulos, Antonios
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912399279259648
author Meng, Chutong
Anastasopoulos, Antonios
author_facet Meng, Chutong
Anastasopoulos, Antonios
contents This paper describes the GMU systems for the IWSLT 2025 low-resource speech translation shared task. We trained systems for all language pairs, except for Levantine Arabic. We fine-tuned SeamlessM4T-v2 for automatic speech recognition (ASR), machine translation (MT), and end-to-end speech translation (E2E ST). The ASR and MT models are also used to form cascaded ST systems. Additionally, we explored various training paradigms for E2E ST fine-tuning, including direct E2E fine-tuning, multi-task training, and parameter initialization using components from fine-tuned ASR and/or MT models. Our results show that (1) direct E2E fine-tuning yields strong results; (2) initializing with a fine-tuned ASR encoder improves ST performance on languages SeamlessM4T-v2 has not been trained on; (3) multi-task training can be slightly helpful.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task
Meng, Chutong
Anastasopoulos, Antonios
Computation and Language
This paper describes the GMU systems for the IWSLT 2025 low-resource speech translation shared task. We trained systems for all language pairs, except for Levantine Arabic. We fine-tuned SeamlessM4T-v2 for automatic speech recognition (ASR), machine translation (MT), and end-to-end speech translation (E2E ST). The ASR and MT models are also used to form cascaded ST systems. Additionally, we explored various training paradigms for E2E ST fine-tuning, including direct E2E fine-tuning, multi-task training, and parameter initialization using components from fine-tuned ASR and/or MT models. Our results show that (1) direct E2E fine-tuning yields strong results; (2) initializing with a fine-tuned ASR encoder improves ST performance on languages SeamlessM4T-v2 has not been trained on; (3) multi-task training can be slightly helpful.
title GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task
topic Computation and Language
url https://arxiv.org/abs/2505.21781