AfroScope: A Framework for Studying the Linguistic Landscape of Africa

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kwon, Sang Yun, Elmadany, AbdelRahim, Abdul-Mageed, Muhammad
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912857133678592
author Kwon, Sang Yun
Elmadany, AbdelRahim
Abdul-Mageed, Muhammad
author_facet Kwon, Sang Yun
Elmadany, AbdelRahim
Abdul-Mageed, Muhammad
contents Language Identification (LID) is the task of determining the language of a given text and is a fundamental preprocessing step that affects the reliability of downstream NLP applications. While recent work has expanded LID coverage for African languages, existing approaches remain limited in (i) the number of supported languages and (ii) their ability to make fine-grained distinctions among closely related varieties. We introduce AfroScope, a unified framework for African LID that includes AfroScope-Data, a dataset covering 713 African languages, and AfroScope-Models, a suite of strong LID models with broad language coverage. To better distinguish highly confusable languages, we propose a hierarchical classification approach that leverages Mirror-Serengeti, a specialized embedding model targeting 29 closely related or geographically proximate languages. This approach improves macro F1 by 4.55 on this confusable subset compared to our best base model. Finally, we analyze cross linguistic transfer and domain effects, offering guidance for building robust African LID systems. We position African LID as an enabling technology for large scale measurement of Africas linguistic landscape in digital text and release AfroScope-Data and AfroScope-Models publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13346
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AfroScope: A Framework for Studying the Linguistic Landscape of Africa
Kwon, Sang Yun
Elmadany, AbdelRahim
Abdul-Mageed, Muhammad
Computation and Language
Language Identification (LID) is the task of determining the language of a given text and is a fundamental preprocessing step that affects the reliability of downstream NLP applications. While recent work has expanded LID coverage for African languages, existing approaches remain limited in (i) the number of supported languages and (ii) their ability to make fine-grained distinctions among closely related varieties. We introduce AfroScope, a unified framework for African LID that includes AfroScope-Data, a dataset covering 713 African languages, and AfroScope-Models, a suite of strong LID models with broad language coverage. To better distinguish highly confusable languages, we propose a hierarchical classification approach that leverages Mirror-Serengeti, a specialized embedding model targeting 29 closely related or geographically proximate languages. This approach improves macro F1 by 4.55 on this confusable subset compared to our best base model. Finally, we analyze cross linguistic transfer and domain effects, offering guidance for building robust African LID systems. We position African LID as an enabling technology for large scale measurement of Africas linguistic landscape in digital text and release AfroScope-Data and AfroScope-Models publicly.
title AfroScope: A Framework for Studying the Linguistic Landscape of Africa
topic Computation and Language
url https://arxiv.org/abs/2601.13346