NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ghimire, Rupak Raj, Subedi, Bipesh, Prasain, Balaram, Poudyal, Prakash, Acharya, Praveen, Karki, Nischal, Tiwari, Rupak, Sharma, Rishikesh Kumar, Poudel, Jenny, Bal, Bal Krishna
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912966254788608
author Ghimire, Rupak Raj
Subedi, Bipesh
Prasain, Balaram
Poudyal, Prakash
Acharya, Praveen
Karki, Nischal
Tiwari, Rupak
Sharma, Rishikesh Kumar
Poudel, Jenny
Bal, Bal Krishna
author_facet Ghimire, Rupak Raj
Subedi, Bipesh
Prasain, Balaram
Poudyal, Prakash
Acharya, Praveen
Karki, Nischal
Tiwari, Rupak
Sharma, Rishikesh Kumar
Poudel, Jenny
Bal, Bal Krishna
contents Modern Translation Systems heavily rely on high-quality, large parallel datasets for state-of-the-art performance. However, such resources are largely unavailable for most of the South Asian languages. Among them, Nepali and Tamang fall into such category, with Tamang being among the least digitally resourced languages in the region. This work addresses the gap by developing NepTam20K, a 20K gold standard parallel corpus, and NepTam80K, an 80K synthetic Nepali-Tamang parallel corpus, both sentence-aligned and designed to support machine translation. The datasets were created through a pipeline involving data scraping from Nepali news and online sources, pre-processing, semantic filtering, balancing for tense and polarity (in NepTam20K dataset), expert translation into Tamang by native speakers of the language, and verification by an expert Tamang linguist. The dataset covers five domains: Agriculture, Health, Education and Technology, Culture, and General Communication. To evaluate the dataset, baseline machine translation experiments were carried out using various multilingual pre-trained models: mBART, M2M-100, NLLB-200, and a vanilla Transformer model. The fine-tuning on the NLLB-200 achieved the highest sacreBLEU scores of 40.92 (Nepali-Tamang) and 45.26 (Tamang-Nepali).
format Preprint
id arxiv_https___arxiv_org_abs_2603_14053
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
Ghimire, Rupak Raj
Subedi, Bipesh
Prasain, Balaram
Poudyal, Prakash
Acharya, Praveen
Karki, Nischal
Tiwari, Rupak
Sharma, Rishikesh Kumar
Poudel, Jenny
Bal, Bal Krishna
Computation and Language
Artificial Intelligence
Machine Learning
Modern Translation Systems heavily rely on high-quality, large parallel datasets for state-of-the-art performance. However, such resources are largely unavailable for most of the South Asian languages. Among them, Nepali and Tamang fall into such category, with Tamang being among the least digitally resourced languages in the region. This work addresses the gap by developing NepTam20K, a 20K gold standard parallel corpus, and NepTam80K, an 80K synthetic Nepali-Tamang parallel corpus, both sentence-aligned and designed to support machine translation. The datasets were created through a pipeline involving data scraping from Nepali news and online sources, pre-processing, semantic filtering, balancing for tense and polarity (in NepTam20K dataset), expert translation into Tamang by native speakers of the language, and verification by an expert Tamang linguist. The dataset covers five domains: Agriculture, Health, Education and Technology, Culture, and General Communication. To evaluate the dataset, baseline machine translation experiments were carried out using various multilingual pre-trained models: mBART, M2M-100, NLLB-200, and a vanilla Transformer model. The fine-tuning on the NLLB-200 achieved the highest sacreBLEU scores of 40.92 (Nepali-Tamang) and 45.26 (Tamang-Nepali).
title NepTam: A Nepali-Tamang Parallel Corpus and Baseline Machine Translation Experiments
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.14053