Graphemic Normalization of the Perso-Arabic Script

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Doctor, Raiomond, Gutkin, Alexander, Johny, Cibu, Roark, Brian, Sproat, Richard
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912355405791232
author Doctor, Raiomond
Gutkin, Alexander
Johny, Cibu
Roark, Brian
Sproat, Richard
author_facet Doctor, Raiomond
Gutkin, Alexander
Johny, Cibu
Roark, Brian
Sproat, Richard
contents Since its original appearance in 1991, the Perso-Arabic script representation in Unicode has grown from 169 to over 440 atomic isolated characters spread over several code pages representing standard letters, various diacritics and punctuation for the original Arabic and numerous other regional orthographic traditions. This paper documents the challenges that Perso-Arabic presents beyond the best-documented languages, such as Arabic and Persian, building on earlier work by the expert community. We particularly focus on the situation in natural language processing (NLP), which is affected by multiple, often neglected, issues such as the use of visually ambiguous yet canonically nonequivalent letters and the mixing of letters from different orthographies. Among the contributing conflating factors are the lack of input methods, the instability of modern orthographies, insufficient literacy, and loss or lack of orthographic tradition. We evaluate the effects of script normalization on eight languages from diverse language families in the Perso-Arabic script diaspora on machine translation and statistical language modeling tasks. Our results indicate statistically significant improvements in performance in most conditions for all the languages considered when normalization is applied. We argue that better understanding and representation of Perso-Arabic script variation within regional orthographic traditions, where those are present, is crucial for further progress of modern computational NLP techniques especially for languages with a paucity of resources.
format Preprint
id arxiv_https___arxiv_org_abs_2210_12273
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Graphemic Normalization of the Perso-Arabic Script
Doctor, Raiomond
Gutkin, Alexander
Johny, Cibu
Roark, Brian
Sproat, Richard
Computation and Language
I.2.7; I.7.2; I.7.1
Since its original appearance in 1991, the Perso-Arabic script representation in Unicode has grown from 169 to over 440 atomic isolated characters spread over several code pages representing standard letters, various diacritics and punctuation for the original Arabic and numerous other regional orthographic traditions. This paper documents the challenges that Perso-Arabic presents beyond the best-documented languages, such as Arabic and Persian, building on earlier work by the expert community. We particularly focus on the situation in natural language processing (NLP), which is affected by multiple, often neglected, issues such as the use of visually ambiguous yet canonically nonequivalent letters and the mixing of letters from different orthographies. Among the contributing conflating factors are the lack of input methods, the instability of modern orthographies, insufficient literacy, and loss or lack of orthographic tradition. We evaluate the effects of script normalization on eight languages from diverse language families in the Perso-Arabic script diaspora on machine translation and statistical language modeling tasks. Our results indicate statistically significant improvements in performance in most conditions for all the languages considered when normalization is applied. We argue that better understanding and representation of Perso-Arabic script variation within regional orthographic traditions, where those are present, is crucial for further progress of modern computational NLP techniques especially for languages with a paucity of resources.
title Graphemic Normalization of the Perso-Arabic Script
topic Computation and Language
I.2.7; I.7.2; I.7.1
url https://arxiv.org/abs/2210.12273