Is It Navajo? Accurate Language Detection in Endangered Athabaskan Languages

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Ivory, Ma, Weicheng, Zhang, Chunhui, Vosoughi, Soroush
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912228039458816
author Yang, Ivory
Ma, Weicheng
Zhang, Chunhui
Vosoughi, Soroush
author_facet Yang, Ivory
Ma, Weicheng
Zhang, Chunhui
Vosoughi, Soroush
contents Endangered languages, such as Navajo - the most widely spoken Native American language - are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. This study evaluates Google's Language Identification (LangID) tool, which does not currently support any Native American languages. To address this, we introduce a random forest classifier trained on Navajo and twenty erroneously suggested languages by LangID. Despite its simplicity, the classifier achieves near-perfect accuracy (97-100%). Additionally, the model demonstrates robustness across other Athabaskan languages - a family of Native American languages spoken primarily in Alaska, the Pacific Northwest, and parts of the Southwestern United States - suggesting its potential for broader application. Our findings underscore the pressing need for NLP systems that prioritize linguistic diversity and adaptability over centralized, one-size-fits-all solutions, especially in supporting underrepresented languages in a multicultural world. This work directly contributes to ongoing efforts to address cultural biases in language models and advocates for the development of culturally localized NLP tools that serve diverse linguistic communities.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15773
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is It Navajo? Accurate Language Detection in Endangered Athabaskan Languages
Yang, Ivory
Ma, Weicheng
Zhang, Chunhui
Vosoughi, Soroush
Computation and Language
Endangered languages, such as Navajo - the most widely spoken Native American language - are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. This study evaluates Google's Language Identification (LangID) tool, which does not currently support any Native American languages. To address this, we introduce a random forest classifier trained on Navajo and twenty erroneously suggested languages by LangID. Despite its simplicity, the classifier achieves near-perfect accuracy (97-100%). Additionally, the model demonstrates robustness across other Athabaskan languages - a family of Native American languages spoken primarily in Alaska, the Pacific Northwest, and parts of the Southwestern United States - suggesting its potential for broader application. Our findings underscore the pressing need for NLP systems that prioritize linguistic diversity and adaptability over centralized, one-size-fits-all solutions, especially in supporting underrepresented languages in a multicultural world. This work directly contributes to ongoing efforts to address cultural biases in language models and advocates for the development of culturally localized NLP tools that serve diverse linguistic communities.
title Is It Navajo? Accurate Language Detection in Endangered Athabaskan Languages
topic Computation and Language
url https://arxiv.org/abs/2501.15773