English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mohammadamini, Mohammad, Jaff, Daban Q., Crego, Josep, Tahon, Marie, Laurent, Antoine
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915905231912960
author Mohammadamini, Mohammad
Jaff, Daban Q.
Crego, Josep
Tahon, Marie
Laurent, Antoine
author_facet Mohammadamini, Mohammad
Jaff, Daban Q.
Crego, Josep
Tahon, Marie
Laurent, Antoine
contents We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40 million Central Kurdish tokens. We evaluate KUTED on the S2TT task and find that orthographic variation significantly degrades Kurdish translation performance, producing nonstandard outputs. To address this, we propose a systematic text standardization approach that yields substantial performance gains and more consistent translations. On a test set separated from TED talks, a fine-tuned Seamless model achieves 15.18 BLEU, and we improve Seamless baseline by 3.0 BLEU on the FLEURS benchmark. We also train a Transformer model from scratch and evaluate a cascaded system that combines Seamless (ASR) with NLLB (MT).
format Preprint
id arxiv_https___arxiv_org_abs_2604_00613
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization
Mohammadamini, Mohammad
Jaff, Daban Q.
Crego, Josep
Tahon, Marie
Laurent, Antoine
Computation and Language
We present KUTED, a speech-to-text translation (S2TT) dataset for Central Kurdish, derived from TED and TEDx talks. The corpus comprises 91,000 sentence pairs, including 170 hours of English audio, 1.65 million English tokens, and 1.40 million Central Kurdish tokens. We evaluate KUTED on the S2TT task and find that orthographic variation significantly degrades Kurdish translation performance, producing nonstandard outputs. To address this, we propose a systematic text standardization approach that yields substantial performance gains and more consistent translations. On a test set separated from TED talks, a fine-tuned Seamless model achieves 15.18 BLEU, and we improve Seamless baseline by 3.0 BLEU on the FLEURS benchmark. We also train a Transformer model from scratch and evaluate a cascaded system that combines Seamless (ASR) with NLLB (MT).
title English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization
topic Computation and Language
url https://arxiv.org/abs/2604.00613