SPRING Lab IITM's submission to Low Resource Indic Language Translation Shared Task

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sayed, Hamees, Joglekar, Advait, Umesh, Srinivasan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912112837656576
author Sayed, Hamees
Joglekar, Advait
Umesh, Srinivasan
author_facet Sayed, Hamees
Joglekar, Advait
Umesh, Srinivasan
contents We develop a robust translation model for four low-resource Indic languages: Khasi, Mizo, Manipuri, and Assamese. Our approach includes a comprehensive pipeline from data collection and preprocessing to training and evaluation, leveraging data from WMT task datasets, BPCC, PMIndia, and OpenLanguageData. To address the scarcity of bilingual data, we use back-translation techniques on monolingual datasets for Mizo and Khasi, significantly expanding our training corpus. We fine-tune the pre-trained NLLB 3.3B model for Assamese, Mizo, and Manipuri, achieving improved performance over the baseline. For Khasi, which is not supported by the NLLB model, we introduce special tokens and train the model on our Khasi corpus. Our training involves masked language modelling, followed by fine-tuning for English-to-Indic and Indic-to-English translations.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00727
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SPRING Lab IITM's submission to Low Resource Indic Language Translation Shared Task
Sayed, Hamees
Joglekar, Advait
Umesh, Srinivasan
Computation and Language
Artificial Intelligence
We develop a robust translation model for four low-resource Indic languages: Khasi, Mizo, Manipuri, and Assamese. Our approach includes a comprehensive pipeline from data collection and preprocessing to training and evaluation, leveraging data from WMT task datasets, BPCC, PMIndia, and OpenLanguageData. To address the scarcity of bilingual data, we use back-translation techniques on monolingual datasets for Mizo and Khasi, significantly expanding our training corpus. We fine-tune the pre-trained NLLB 3.3B model for Assamese, Mizo, and Manipuri, achieving improved performance over the baseline. For Khasi, which is not supported by the NLLB model, we introduce special tokens and train the model on our Khasi corpus. Our training involves masked language modelling, followed by fine-tuning for English-to-Indic and Indic-to-English translations.
title SPRING Lab IITM's submission to Low Resource Indic Language Translation Shared Task
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2411.00727