QuranMorph: Morphologically Annotated Quranic Corpus

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Akra, Diyam, Hammouda, Tymaa, Jarrar, Mustafa
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908416646053888
author Akra, Diyam
Hammouda, Tymaa
Jarrar, Mustafa
author_facet Akra, Diyam
Hammouda, Tymaa
Jarrar, Mustafa
contents We present the QuranMorph corpus, a morphologically annotated corpus for the Quran (77,429 tokens). Each token in the QuranMorph was manually lemmatized and tagged with its part-of-speech by three expert linguists. The lemmatization process utilized lemmas from Qabas, an Arabic lexicographic database linked with 110 lexicons and corpora of 2 million tokens. The part-of-speech tagging was performed using the fine-grained SAMA/Qabas tagset, which encompasses 40 tags. As shown in this paper, this rich lemmatization and POS tagset enabled the QuranMorph corpus to be inter-linked with many linguistic resources. The corpus is open-source and publicly available as part of the SinaLab resources at (https://sina.birzeit.edu/quran)
format Preprint
id arxiv_https___arxiv_org_abs_2506_18148
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QuranMorph: Morphologically Annotated Quranic Corpus
Akra, Diyam
Hammouda, Tymaa
Jarrar, Mustafa
Computation and Language
Artificial Intelligence
We present the QuranMorph corpus, a morphologically annotated corpus for the Quran (77,429 tokens). Each token in the QuranMorph was manually lemmatized and tagged with its part-of-speech by three expert linguists. The lemmatization process utilized lemmas from Qabas, an Arabic lexicographic database linked with 110 lexicons and corpora of 2 million tokens. The part-of-speech tagging was performed using the fine-grained SAMA/Qabas tagset, which encompasses 40 tags. As shown in this paper, this rich lemmatization and POS tagset enabled the QuranMorph corpus to be inter-linked with many linguistic resources. The corpus is open-source and publicly available as part of the SinaLab resources at (https://sina.birzeit.edu/quran)
title QuranMorph: Morphologically Annotated Quranic Corpus
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.18148