BhashaSetu: A Data-Centric Approach to Low-Resource Machine Translation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thakkar, Param, Yadav, Anushka, Tiemann, Michael, Mehta, Abhi, Bhasin, Akshita, Khedkar, Shrinivas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914604845629440
author Thakkar, Param
Yadav, Anushka
Tiemann, Michael
Mehta, Abhi
Bhasin, Akshita
Khedkar, Shrinivas
author_facet Thakkar, Param
Yadav, Anushka
Tiemann, Michael
Mehta, Abhi
Bhasin, Akshita
Khedkar, Shrinivas
contents We present BhashaSetu, a linguistically enriched English--Marathi parallel dataset addressing persistent data limitations in low-resource neural machine translation (NMT). Marathi, spoken by over 95 million people, remains underrepresented in high-quality parallel corpora across diverse domains. Our dataset comprises 2.78 million sentence pairs from heterogeneous sources including news, politics, healthcare, literature, and culture, with stemmed and lemmatized representations to support morphology-aware analysis. We benchmark multiple state-of-the-art translation models using BLEU, spBLEU, chrF++, and TER metrics, and conduct parameter-efficient fine-tuning of NLLB-200-distilled-600M using LoRA. A key finding from our ablation: corpus-level deduplication is the single largest preprocessing contributor to downstream quality (removing it reduces performance by 1.17 BLEU and 2.21 chrF++), demonstrating that disciplined cross-source corpus hygiene is a low-cost, high-impact intervention for low-resource, morphologically rich languages. The dataset is publicly released to promote reproducible and linguistically informed low-resource NMT research.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27050
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BhashaSetu: A Data-Centric Approach to Low-Resource Machine Translation
Thakkar, Param
Yadav, Anushka
Tiemann, Michael
Mehta, Abhi
Bhasin, Akshita
Khedkar, Shrinivas
Computation and Language
Machine Learning
We present BhashaSetu, a linguistically enriched English--Marathi parallel dataset addressing persistent data limitations in low-resource neural machine translation (NMT). Marathi, spoken by over 95 million people, remains underrepresented in high-quality parallel corpora across diverse domains. Our dataset comprises 2.78 million sentence pairs from heterogeneous sources including news, politics, healthcare, literature, and culture, with stemmed and lemmatized representations to support morphology-aware analysis. We benchmark multiple state-of-the-art translation models using BLEU, spBLEU, chrF++, and TER metrics, and conduct parameter-efficient fine-tuning of NLLB-200-distilled-600M using LoRA. A key finding from our ablation: corpus-level deduplication is the single largest preprocessing contributor to downstream quality (removing it reduces performance by 1.17 BLEU and 2.21 chrF++), demonstrating that disciplined cross-source corpus hygiene is a low-cost, high-impact intervention for low-resource, morphologically rich languages. The dataset is publicly released to promote reproducible and linguistically informed low-resource NMT research.
title BhashaSetu: A Data-Centric Approach to Low-Resource Machine Translation
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2605.27050