Dialectal and Low-Resource Machine Translation for Aromanian

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jerpelea, Alexandru-Iulius, Rădoi, Alina, Nisioi, Sergiu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909450310254592
author Jerpelea, Alexandru-Iulius
Rădoi, Alina
Nisioi, Sergiu
author_facet Jerpelea, Alexandru-Iulius
Rădoi, Alina
Nisioi, Sergiu
contents This paper presents the process of building a neural machine translation system with support for English, Romanian, and Aromanian - an endangered Eastern Romance language. The primary contribution of this research is twofold: (1) the creation of the most extensive Aromanian-Romanian parallel corpus to date, consisting of 79,000 sentence pairs, and (2) the development and comparative analysis of several machine translation models optimized for Aromanian. To accomplish this, we introduce a suite of auxiliary tools, including a language-agnostic sentence embedding model for text mining and automated evaluation, complemented by a diacritics conversion system for different writing standards. This research brings contributions to both computational linguistics and language preservation efforts by establishing essential resources for a historically under-resourced language. All datasets, trained models, and associated tools are public: https://huggingface.co/aronlp and https://arotranslate.com
format Preprint
id arxiv_https___arxiv_org_abs_2410_17728
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dialectal and Low-Resource Machine Translation for Aromanian
Jerpelea, Alexandru-Iulius
Rădoi, Alina
Nisioi, Sergiu
Computation and Language
This paper presents the process of building a neural machine translation system with support for English, Romanian, and Aromanian - an endangered Eastern Romance language. The primary contribution of this research is twofold: (1) the creation of the most extensive Aromanian-Romanian parallel corpus to date, consisting of 79,000 sentence pairs, and (2) the development and comparative analysis of several machine translation models optimized for Aromanian. To accomplish this, we introduce a suite of auxiliary tools, including a language-agnostic sentence embedding model for text mining and automated evaluation, complemented by a diacritics conversion system for different writing standards. This research brings contributions to both computational linguistics and language preservation efforts by establishing essential resources for a historically under-resourced language. All datasets, trained models, and associated tools are public: https://huggingface.co/aronlp and https://arotranslate.com
title Dialectal and Low-Resource Machine Translation for Aromanian
topic Computation and Language
url https://arxiv.org/abs/2410.17728