Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dhawan, Aashish, Driggers-Ellis, Christopher, Grant, Christan, Wang, Daisy Zhe
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911699269844992
author Dhawan, Aashish
Driggers-Ellis, Christopher
Grant, Christan
Wang, Daisy Zhe
author_facet Dhawan, Aashish
Driggers-Ellis, Christopher
Grant, Christan
Wang, Daisy Zhe
contents Low-resource indigenous languages often lack the parallel corpora required for effective neural machine translation (NMT). Synthetic data generation offers a practical strategy for mitigating this limitation in data-scarce settings. In this work, we augment curated parallel datasets for indigenous languages of the Americas with synthetic sentence pairs generated using a high-capacity multilingual translation model. We fine-tune a multilingual mBART model on curated-only and synthetically augmented data and evaluate translation quality using chrF++, the primary metric used in recent AmericasNLP shared tasks for agglutinative languages. We further apply language-specific preprocessing, including orthographic normalization and noise-aware filtering, to reduce corpus artifacts. Experiments on Guarani-Spanish and Quechua-Spanish translation show consistent chrF++ improvements from synthetic data augmentation, while diagnostic experiments on Aymara highlight the limitations of generic preprocessing for highly agglutinative languages.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03135
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing
Dhawan, Aashish
Driggers-Ellis, Christopher
Grant, Christan
Wang, Daisy Zhe
Computation and Language
Low-resource indigenous languages often lack the parallel corpora required for effective neural machine translation (NMT). Synthetic data generation offers a practical strategy for mitigating this limitation in data-scarce settings. In this work, we augment curated parallel datasets for indigenous languages of the Americas with synthetic sentence pairs generated using a high-capacity multilingual translation model. We fine-tune a multilingual mBART model on curated-only and synthetically augmented data and evaluate translation quality using chrF++, the primary metric used in recent AmericasNLP shared tasks for agglutinative languages. We further apply language-specific preprocessing, including orthographic normalization and noise-aware filtering, to reduce corpus artifacts. Experiments on Guarani-Spanish and Quechua-Spanish translation show consistent chrF++ improvements from synthetic data augmentation, while diagnostic experiments on Aymara highlight the limitations of generic preprocessing for highly agglutinative languages.
title Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing
topic Computation and Language
url https://arxiv.org/abs/2601.03135