EDU-NER-2025: Named Entity Recognition in Urdu Educational Texts using XLM-RoBERTa with X (formerly Twitter)

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ullah, Fida, Ahmad, Muhammad, Zamir, Muhammad Tayyab, Arif, Muhammad, sidorov, Grigori, Riverón, Edgardo Manuel Felipe, Gelbukh, Alexander
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912346459340800
author Ullah, Fida
Ahmad, Muhammad
Zamir, Muhammad Tayyab
Arif, Muhammad
sidorov, Grigori
Riverón, Edgardo Manuel Felipe
Gelbukh, Alexander
author_facet Ullah, Fida
Ahmad, Muhammad
Zamir, Muhammad Tayyab
Arif, Muhammad
sidorov, Grigori
Riverón, Edgardo Manuel Felipe
Gelbukh, Alexander
contents Named Entity Recognition (NER) plays a pivotal role in various Natural Language Processing (NLP) tasks by identifying and classifying named entities (NEs) from unstructured data into predefined categories such as person, organization, location, date, and time. While extensive research exists for high-resource languages and general domains, NER in Urdu particularly within domain-specific contexts like education remains significantly underexplored. This is Due to lack of annotated datasets for educational content which limits the ability of existing models to accurately identify entities such as academic roles, course names, and institutional terms, underscoring the urgent need for targeted resources in this domain. To the best of our knowledge, no dataset exists in the domain of the Urdu language for this purpose. To achieve this objective this study makes three key contributions. Firstly, we created a manually annotated dataset in the education domain, named EDU-NER-2025, which contains 13 unique most important entities related to education domain. Second, we describe our annotation process and guidelines in detail and discuss the challenges of labelling EDU-NER-2025 dataset. Third, we addressed and analyzed key linguistic challenges, such as morphological complexity and ambiguity, which are prevalent in formal Urdu texts.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18142
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EDU-NER-2025: Named Entity Recognition in Urdu Educational Texts using XLM-RoBERTa with X (formerly Twitter)
Ullah, Fida
Ahmad, Muhammad
Zamir, Muhammad Tayyab
Arif, Muhammad
sidorov, Grigori
Riverón, Edgardo Manuel Felipe
Gelbukh, Alexander
Computation and Language
Artificial Intelligence
Named Entity Recognition (NER) plays a pivotal role in various Natural Language Processing (NLP) tasks by identifying and classifying named entities (NEs) from unstructured data into predefined categories such as person, organization, location, date, and time. While extensive research exists for high-resource languages and general domains, NER in Urdu particularly within domain-specific contexts like education remains significantly underexplored. This is Due to lack of annotated datasets for educational content which limits the ability of existing models to accurately identify entities such as academic roles, course names, and institutional terms, underscoring the urgent need for targeted resources in this domain. To the best of our knowledge, no dataset exists in the domain of the Urdu language for this purpose. To achieve this objective this study makes three key contributions. Firstly, we created a manually annotated dataset in the education domain, named EDU-NER-2025, which contains 13 unique most important entities related to education domain. Second, we describe our annotation process and guidelines in detail and discuss the challenges of labelling EDU-NER-2025 dataset. Third, we addressed and analyzed key linguistic challenges, such as morphological complexity and ambiguity, which are prevalent in formal Urdu texts.
title EDU-NER-2025: Named Entity Recognition in Urdu Educational Texts using XLM-RoBERTa with X (formerly Twitter)
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.18142