Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: De Mel, Yomal, Wickramasinghe, Kasun, de Silva, Nisansa, Ranathunga, Surangika
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910856670871552
author De Mel, Yomal
Wickramasinghe, Kasun
de Silva, Nisansa
Ranathunga, Surangika
author_facet De Mel, Yomal
Wickramasinghe, Kasun
de Silva, Nisansa
Ranathunga, Surangika
contents Due to reasons of convenience and lack of tech literacy, transliteration (i.e., Romanizing native scripts instead of using localization tools) is eminently prevalent in the context of low-resource languages such as Sinhala, which have their own writing script. In this study, our focus is on Romanized Sinhala transliteration. We propose two methods to address this problem: Our baseline is a rule-based method, which is then compared against our second method where we approach the transliteration problem as a sequence-to-sequence task akin to the established Neural Machine Translation (NMT) task. For the latter, we propose a Transformer-based Encode-Decoder solution. We witnessed that the Transformer-based method could grab many ad-hoc patterns within the Romanized scripts compared to the rule-based method. The code base associated with this paper is available on GitHub - https://github.com/kasunw22/Sinhala-Transliterator/
format Preprint
id arxiv_https___arxiv_org_abs_2501_00529
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches
De Mel, Yomal
Wickramasinghe, Kasun
de Silva, Nisansa
Ranathunga, Surangika
Computation and Language
F.2.2, I.2.7 68T50
Due to reasons of convenience and lack of tech literacy, transliteration (i.e., Romanizing native scripts instead of using localization tools) is eminently prevalent in the context of low-resource languages such as Sinhala, which have their own writing script. In this study, our focus is on Romanized Sinhala transliteration. We propose two methods to address this problem: Our baseline is a rule-based method, which is then compared against our second method where we approach the transliteration problem as a sequence-to-sequence task akin to the established Neural Machine Translation (NMT) task. For the latter, we propose a Transformer-based Encode-Decoder solution. We witnessed that the Transformer-based method could grab many ad-hoc patterns within the Romanized scripts compared to the rule-based method. The code base associated with this paper is available on GitHub - https://github.com/kasunw22/Sinhala-Transliterator/
title Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches
topic Computation and Language
F.2.2, I.2.7 68T50
url https://arxiv.org/abs/2501.00529