Script-Agnostic Language Identification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Agarwal, Milind, Otten, Joshua, Anastasopoulos, Antonios
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911933379117056
author Agarwal, Milind
Otten, Joshua
Anastasopoulos, Antonios
author_facet Agarwal, Milind
Otten, Joshua
Anastasopoulos, Antonios
contents Language identification is used as the first step in many data collection and crawling efforts because it allows us to sort online text into language-specific buckets. However, many modern languages, such as Konkani, Kashmiri, Punjabi etc., are synchronically written in several scripts. Moreover, languages with different writing systems do not share significant lexical, semantic, and syntactic properties in neural representation spaces, which is a disadvantage for closely related languages and low-resource languages, especially those from the Indian Subcontinent. To counter this, we propose learning script-agnostic representations using several different experimental strategies (upscaling, flattening, and script mixing) focusing on four major Dravidian languages (Tamil, Telugu, Kannada, and Malayalam). We find that word-level script randomization and exposure to a language written in multiple scripts is extremely valuable for downstream script-agnostic language identification, while also maintaining competitive performance on naturally occurring text.
format Preprint
id arxiv_https___arxiv_org_abs_2406_17901
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Script-Agnostic Language Identification
Agarwal, Milind
Otten, Joshua
Anastasopoulos, Antonios
Computation and Language
Language identification is used as the first step in many data collection and crawling efforts because it allows us to sort online text into language-specific buckets. However, many modern languages, such as Konkani, Kashmiri, Punjabi etc., are synchronically written in several scripts. Moreover, languages with different writing systems do not share significant lexical, semantic, and syntactic properties in neural representation spaces, which is a disadvantage for closely related languages and low-resource languages, especially those from the Indian Subcontinent. To counter this, we propose learning script-agnostic representations using several different experimental strategies (upscaling, flattening, and script mixing) focusing on four major Dravidian languages (Tamil, Telugu, Kannada, and Malayalam). We find that word-level script randomization and exposure to a language written in multiple scripts is extremely valuable for downstream script-agnostic language identification, while also maintaining competitive performance on naturally occurring text.
title Script-Agnostic Language Identification
topic Computation and Language
url https://arxiv.org/abs/2406.17901