Pivot Language for Low-Resource Machine Translation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Talwar, Abhimanyu, Laasri, Julien
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913850669924352
author Talwar, Abhimanyu
Laasri, Julien
author_facet Talwar, Abhimanyu
Laasri, Julien
contents Certain pairs of languages suffer from lack of a parallel corpus which is large in size and diverse in domain. One of the ways this is overcome is via use of a pivot language. In this paper we use Hindi as a pivot language to translate Nepali into English. We describe what makes Hindi a good candidate for the pivot. We discuss ways in which a pivot language can be used, and use two such approaches - the Transfer Method (fully supervised) and Backtranslation (semi-supervised) - to translate Nepali into English. Using the former, we are able to achieve a devtest Set SacreBLEU score of 14.2, which improves the baseline fully supervised score reported by (Guzman et al., 2019) by 6.6 points. While we are slightly below the semi-supervised baseline score of 15.1, we discuss what may have caused this under-performance, and suggest scope for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14553
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pivot Language for Low-Resource Machine Translation
Talwar, Abhimanyu
Laasri, Julien
Computation and Language
Machine Learning
68T50
I.2.7
Certain pairs of languages suffer from lack of a parallel corpus which is large in size and diverse in domain. One of the ways this is overcome is via use of a pivot language. In this paper we use Hindi as a pivot language to translate Nepali into English. We describe what makes Hindi a good candidate for the pivot. We discuss ways in which a pivot language can be used, and use two such approaches - the Transfer Method (fully supervised) and Backtranslation (semi-supervised) - to translate Nepali into English. Using the former, we are able to achieve a devtest Set SacreBLEU score of 14.2, which improves the baseline fully supervised score reported by (Guzman et al., 2019) by 6.6 points. While we are slightly below the semi-supervised baseline score of 15.1, we discuss what may have caused this under-performance, and suggest scope for future work.
title Pivot Language for Low-Resource Machine Translation
topic Computation and Language
Machine Learning
68T50
I.2.7
url https://arxiv.org/abs/2505.14553