Unicode Normalization and Grapheme Parsing of Indic Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ansary, Nazmuddoha, Adib, Quazi Adibur Rahman, Reasat, Tahsin, Sushmit, Asif Shahriyar, Humayun, Ahmed Imtiaz, Mehnaz, Sazia, Fatema, Kanij, Rashid, Mohammad Mamun Or, Sadeque, Farig
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916260187471872
author Ansary, Nazmuddoha
Adib, Quazi Adibur Rahman
Reasat, Tahsin
Sushmit, Asif Shahriyar
Humayun, Ahmed Imtiaz
Mehnaz, Sazia
Fatema, Kanij
Rashid, Mohammad Mamun Or
Sadeque, Farig
author_facet Ansary, Nazmuddoha
Adib, Quazi Adibur Rahman
Reasat, Tahsin
Sushmit, Asif Shahriyar
Humayun, Ahmed Imtiaz
Mehnaz, Sazia
Fatema, Kanij
Rashid, Mohammad Mamun Or
Sadeque, Farig
contents Writing systems of Indic languages have orthographic syllables, also known as complex graphemes, as unique horizontal units. A prominent feature of these languages is these complex grapheme units that comprise consonants/consonant conjuncts, vowel diacritics, and consonant diacritics, which, together make a unique Language. Unicode-based writing schemes of these languages often disregard this feature of these languages and encode words as linear sequences of Unicode characters using an intricate scheme of connector characters and font interpreters. Due to this way of using a few dozen Unicode glyphs to write thousands of different unique glyphs (complex graphemes), there are serious ambiguities that lead to malformed words. In this paper, we are proposing two libraries: i) a normalizer for normalizing inconsistencies caused by a Unicode-based encoding scheme for Indic languages and ii) a grapheme parser for Abugida text. It deconstructs words into visually distinct orthographic syllables or complex graphemes and their constituents. Our proposed normalizer is a more efficient and effective tool than the previously used IndicNLP normalizer. Moreover, our parser and normalizer are also suitable tools for general Abugida text processing as they performed well in our robust word-based and NLP experiments. We report the pipeline for the scripts of 7 languages in this work and develop the framework for the integration of more scripts.
format Preprint
id arxiv_https___arxiv_org_abs_2306_01743
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Unicode Normalization and Grapheme Parsing of Indic Languages
Ansary, Nazmuddoha
Adib, Quazi Adibur Rahman
Reasat, Tahsin
Sushmit, Asif Shahriyar
Humayun, Ahmed Imtiaz
Mehnaz, Sazia
Fatema, Kanij
Rashid, Mohammad Mamun Or
Sadeque, Farig
Computation and Language
Writing systems of Indic languages have orthographic syllables, also known as complex graphemes, as unique horizontal units. A prominent feature of these languages is these complex grapheme units that comprise consonants/consonant conjuncts, vowel diacritics, and consonant diacritics, which, together make a unique Language. Unicode-based writing schemes of these languages often disregard this feature of these languages and encode words as linear sequences of Unicode characters using an intricate scheme of connector characters and font interpreters. Due to this way of using a few dozen Unicode glyphs to write thousands of different unique glyphs (complex graphemes), there are serious ambiguities that lead to malformed words. In this paper, we are proposing two libraries: i) a normalizer for normalizing inconsistencies caused by a Unicode-based encoding scheme for Indic languages and ii) a grapheme parser for Abugida text. It deconstructs words into visually distinct orthographic syllables or complex graphemes and their constituents. Our proposed normalizer is a more efficient and effective tool than the previously used IndicNLP normalizer. Moreover, our parser and normalizer are also suitable tools for general Abugida text processing as they performed well in our robust word-based and NLP experiments. We report the pipeline for the scripts of 7 languages in this work and develop the framework for the integration of more scripts.
title Unicode Normalization and Grapheme Parsing of Indic Languages
topic Computation and Language
url https://arxiv.org/abs/2306.01743