Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kimera, Richard, Heo, Dongnyeong, Rim, Daniela N., Choi, Heeyoul
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918009248940032
author Kimera, Richard
Heo, Dongnyeong
Rim, Daniela N.
Choi, Heeyoul
author_facet Kimera, Richard
Heo, Dongnyeong
Rim, Daniela N.
Choi, Heeyoul
contents In this paper,we explore the application of Back translation (BT) as a semi-supervised technique to enhance Neural Machine Translation(NMT) models for the English-Luganda language pair, specifically addressing the challenges faced by low-resource languages. The purpose of our study is to demonstrate how BT can mitigate the scarcity of bilingual data by generating synthetic data from monolingual corpora. Our methodology involves developing custom NMT models using both publicly available and web-crawled data, and applying Iterative and Incremental Back translation techniques. We strategically select datasets for incremental back translation across multiple small datasets, which is a novel element of our approach. The results of our study show significant improvements, with translation performance for the English-Luganda pair exceeding previous benchmarks by more than 10 BLEU score units across all translation directions. Additionally, our evaluation incorporates comprehensive assessment metrics such as SacreBLEU, ChrF2, and TER, providing a nuanced understanding of translation quality. The conclusion drawn from our research confirms the efficacy of BT when strategically curated datasets are utilized, establishing new performance benchmarks and demonstrating the potential of BT in enhancing NMT models for low-resource languages.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02463
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda
Kimera, Richard
Heo, Dongnyeong
Rim, Daniela N.
Choi, Heeyoul
Computation and Language
In this paper,we explore the application of Back translation (BT) as a semi-supervised technique to enhance Neural Machine Translation(NMT) models for the English-Luganda language pair, specifically addressing the challenges faced by low-resource languages. The purpose of our study is to demonstrate how BT can mitigate the scarcity of bilingual data by generating synthetic data from monolingual corpora. Our methodology involves developing custom NMT models using both publicly available and web-crawled data, and applying Iterative and Incremental Back translation techniques. We strategically select datasets for incremental back translation across multiple small datasets, which is a novel element of our approach. The results of our study show significant improvements, with translation performance for the English-Luganda pair exceeding previous benchmarks by more than 10 BLEU score units across all translation directions. Additionally, our evaluation incorporates comprehensive assessment metrics such as SacreBLEU, ChrF2, and TER, providing a nuanced understanding of translation quality. The conclusion drawn from our research confirms the efficacy of BT when strategically curated datasets are utilized, establishing new performance benchmarks and demonstrating the potential of BT in enhancing NMT models for low-resource languages.
title Data Augmentation With Back translation for Low Resource languages: A case of English and Luganda
topic Computation and Language
url https://arxiv.org/abs/2505.02463