Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910856670871552 |
|---|---|
| author | De Mel, Yomal Wickramasinghe, Kasun de Silva, Nisansa Ranathunga, Surangika |
| author_facet | De Mel, Yomal Wickramasinghe, Kasun de Silva, Nisansa Ranathunga, Surangika |
| contents | Due to reasons of convenience and lack of tech literacy, transliteration (i.e., Romanizing native scripts instead of using localization tools) is eminently prevalent in the context of low-resource languages such as Sinhala, which have their own writing script. In this study, our focus is on Romanized Sinhala transliteration. We propose two methods to address this problem: Our baseline is a rule-based method, which is then compared against our second method where we approach the transliteration problem as a sequence-to-sequence task akin to the established Neural Machine Translation (NMT) task. For the latter, we propose a Transformer-based Encode-Decoder solution. We witnessed that the Transformer-based method could grab many ad-hoc patterns within the Romanized scripts compared to the rule-based method. The code base associated with this paper is available on GitHub - https://github.com/kasunw22/Sinhala-Transliterator/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_00529 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches De Mel, Yomal Wickramasinghe, Kasun de Silva, Nisansa Ranathunga, Surangika Computation and Language F.2.2, I.2.7 68T50 Due to reasons of convenience and lack of tech literacy, transliteration (i.e., Romanizing native scripts instead of using localization tools) is eminently prevalent in the context of low-resource languages such as Sinhala, which have their own writing script. In this study, our focus is on Romanized Sinhala transliteration. We propose two methods to address this problem: Our baseline is a rule-based method, which is then compared against our second method where we approach the transliteration problem as a sequence-to-sequence task akin to the established Neural Machine Translation (NMT) task. For the latter, we propose a Transformer-based Encode-Decoder solution. We witnessed that the Transformer-based method could grab many ad-hoc patterns within the Romanized scripts compared to the rule-based method. The code base associated with this paper is available on GitHub - https://github.com/kasunw22/Sinhala-Transliterator/ |
| title | Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches |
| topic | Computation and Language F.2.2, I.2.7 68T50 |
| url | https://arxiv.org/abs/2501.00529 |