Konooz: Multi-domain Multi-dialect Corpus for Named Entity Recognition

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hamad, Nagham, Khalilia, Mohammed, Jarrar, Mustafa
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913893954093056
author Hamad, Nagham
Khalilia, Mohammed
Jarrar, Mustafa
author_facet Hamad, Nagham
Khalilia, Mohammed
Jarrar, Mustafa
contents We introduce Konooz, a novel multi-dimensional corpus covering 16 Arabic dialects across 10 domains, resulting in 160 distinct corpora. The corpus comprises about 777k tokens, carefully collected and manually annotated with 21 entity types using both nested and flat annotation schemes - using the Wojood guidelines. While Konooz is useful for various NLP tasks like domain adaptation and transfer learning, this paper primarily focuses on benchmarking existing Arabic Named Entity Recognition (NER) models, especially cross-domain and cross-dialect model performance. Our benchmarking of four Arabic NER models using Konooz reveals a significant drop in performance of up to 38% when compared to the in-distribution data. Furthermore, we present an in-depth analysis of domain and dialect divergence and the impact of resource scarcity. We also measured the overlap between domains and dialects using the Maximum Mean Discrepancy (MMD) metric, and illustrated why certain NER models perform better on specific dialects and domains. Konooz is open-source and publicly available at https://sina.birzeit.edu/wojood/#download
format Preprint
id arxiv_https___arxiv_org_abs_2506_12615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Konooz: Multi-domain Multi-dialect Corpus for Named Entity Recognition
Hamad, Nagham
Khalilia, Mohammed
Jarrar, Mustafa
Computation and Language
Artificial Intelligence
We introduce Konooz, a novel multi-dimensional corpus covering 16 Arabic dialects across 10 domains, resulting in 160 distinct corpora. The corpus comprises about 777k tokens, carefully collected and manually annotated with 21 entity types using both nested and flat annotation schemes - using the Wojood guidelines. While Konooz is useful for various NLP tasks like domain adaptation and transfer learning, this paper primarily focuses on benchmarking existing Arabic Named Entity Recognition (NER) models, especially cross-domain and cross-dialect model performance. Our benchmarking of four Arabic NER models using Konooz reveals a significant drop in performance of up to 38% when compared to the in-distribution data. Furthermore, we present an in-depth analysis of domain and dialect divergence and the impact of resource scarcity. We also measured the overlap between domains and dialects using the Maximum Mean Discrepancy (MMD) metric, and illustrated why certain NER models perform better on specific dialects and domains. Konooz is open-source and publicly available at https://sina.birzeit.edu/wojood/#download
title Konooz: Multi-domain Multi-dialect Corpus for Named Entity Recognition
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.12615