Improving Informally Romanized Language Identification

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Benton, Adrian, Gutkin, Alexander, Kirov, Christo, Roark, Brian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912715341037568
author Benton, Adrian
Gutkin, Alexander
Kirov, Christo
Roark, Brian
author_facet Benton, Adrian
Gutkin, Alexander
Kirov, Christo
Roark, Brian
contents The Latin script is often used to informally write languages with non-Latin native scripts. In many cases (e.g., most languages in India), the lack of conventional spelling in the Latin script results in high spelling variability. Such romanization renders languages that are normally easily distinguished due to being written in different scripts - Hindi and Urdu, for example - highly confusable. In this work, we increase language identification (LID) accuracy for romanized text by improving the methods used to synthesize training sets. We find that training on synthetic samples which incorporate natural spelling variation yields higher LID system accuracy than including available naturally occurring examples in the training set, or even training higher capacity models. We demonstrate new state-of-the-art LID performance on romanized text from 20 Indic languages in the Bhasha-Abhijnaanam evaluation set (Madhani et al., 2023a), improving test F1 from the reported 74.7% (using a pretrained neural model) to 85.4% using a linear classifier trained solely on synthetic data and 88.2% when also training on available harvested text.
format Preprint
id arxiv_https___arxiv_org_abs_2504_21540
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Informally Romanized Language Identification
Benton, Adrian
Gutkin, Alexander
Kirov, Christo
Roark, Brian
Computation and Language
The Latin script is often used to informally write languages with non-Latin native scripts. In many cases (e.g., most languages in India), the lack of conventional spelling in the Latin script results in high spelling variability. Such romanization renders languages that are normally easily distinguished due to being written in different scripts - Hindi and Urdu, for example - highly confusable. In this work, we increase language identification (LID) accuracy for romanized text by improving the methods used to synthesize training sets. We find that training on synthetic samples which incorporate natural spelling variation yields higher LID system accuracy than including available naturally occurring examples in the training set, or even training higher capacity models. We demonstrate new state-of-the-art LID performance on romanized text from 20 Indic languages in the Bhasha-Abhijnaanam evaluation set (Madhani et al., 2023a), improving test F1 from the reported 74.7% (using a pretrained neural model) to 85.4% using a linear classifier trained solely on synthetic data and 88.2% when also training on available harvested text.
title Improving Informally Romanized Language Identification
topic Computation and Language
url https://arxiv.org/abs/2504.21540