Constructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local Languages

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lopo, Joanito Agili, Tanone, Radius
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914737126637568
author Lopo, Joanito Agili
Tanone, Radius
author_facet Lopo, Joanito Agili
Tanone, Radius
contents In Indonesia, local languages play an integral role in the culture. However, the available Indonesian language resources still fall into the category of limited data in the Natural Language Processing (NLP) field. This is become problematic when build NLP model for these languages. To address this gap, we introduce Bhinneka Korpus, a multilingual parallel corpus featuring five Indonesian local languages. Our goal is to enhance access and utilization of these resources, extending their reach within the country. We explained in a detail the dataset collection process and associated challenges. Additionally, we experimented with translation task using the IBM Model 1 due to data constraints. The result showed that the performance of each language already shows good indications for further development. Challenges such as lexical variation, smoothing effects, and cross-linguistic variability are discussed. We intend to evaluate the corpus using advanced NLP techniques for low-resource languages, paving the way for multilingual translation models.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01009
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Constructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local Languages
Lopo, Joanito Agili
Tanone, Radius
Computation and Language
In Indonesia, local languages play an integral role in the culture. However, the available Indonesian language resources still fall into the category of limited data in the Natural Language Processing (NLP) field. This is become problematic when build NLP model for these languages. To address this gap, we introduce Bhinneka Korpus, a multilingual parallel corpus featuring five Indonesian local languages. Our goal is to enhance access and utilization of these resources, extending their reach within the country. We explained in a detail the dataset collection process and associated challenges. Additionally, we experimented with translation task using the IBM Model 1 due to data constraints. The result showed that the performance of each language already shows good indications for further development. Challenges such as lexical variation, smoothing effects, and cross-linguistic variability are discussed. We intend to evaluate the corpus using advanced NLP techniques for low-resource languages, paving the way for multilingual translation models.
title Constructing and Expanding Low-Resource and Underrepresented Parallel Datasets for Indonesian Local Languages
topic Computation and Language
url https://arxiv.org/abs/2404.01009