Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Micallef, Kurt, Habash, Nizar, Borg, Claudia
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909899033673728
author Micallef, Kurt
Habash, Nizar
Borg, Claudia
author_facet Micallef, Kurt
Habash, Nizar
Borg, Claudia
contents Maltese is a unique Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. Despite its Semitic roots, its orthography is based on the Latin script, creating a gap between it and its closest linguistic relatives in Arabic. In this paper, we explore whether Arabic-language resources can support Maltese natural language processing (NLP) through cross-lingual augmentation techniques. We investigate multiple strategies for aligning Arabic textual data with Maltese, including various transliteration schemes and machine translation (MT) approaches. As part of this, we also introduce novel transliteration systems that better represent Maltese orthography. We evaluate the impact of these augmentations on monolingual and mutlilingual models and demonstrate that Arabic-based augmentation can significantly benefit Maltese NLP tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_12853
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data
Micallef, Kurt
Habash, Nizar
Borg, Claudia
Computation and Language
Maltese is a unique Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. Despite its Semitic roots, its orthography is based on the Latin script, creating a gap between it and its closest linguistic relatives in Arabic. In this paper, we explore whether Arabic-language resources can support Maltese natural language processing (NLP) through cross-lingual augmentation techniques. We investigate multiple strategies for aligning Arabic textual data with Maltese, including various transliteration schemes and machine translation (MT) approaches. As part of this, we also introduce novel transliteration systems that better represent Maltese orthography. We evaluate the impact of these augmentations on monolingual and mutlilingual models and demonstrate that Arabic-based augmentation can significantly benefit Maltese NLP tasks.
title Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data
topic Computation and Language
url https://arxiv.org/abs/2509.12853