KuBERT: Central Kurdish BERT Model and Its Application for Sentiment Analysis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Awlla, Kozhin muhealddin, Veisi, Hadi, Abdullah, Abdulhady Abas
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911165975625728
author Awlla, Kozhin muhealddin
Veisi, Hadi
Abdullah, Abdulhady Abas
author_facet Awlla, Kozhin muhealddin
Veisi, Hadi
Abdullah, Abdulhady Abas
contents This paper enhances the study of sentiment analysis for the Central Kurdish language by integrating the Bidirectional Encoder Representations from Transformers (BERT) into Natural Language Processing techniques. Kurdish is a low-resourced language, having a high level of linguistic diversity with minimal computational resources, making sentiment analysis somewhat challenging. Earlier, this was done using a traditional word embedding model, such as Word2Vec, but with the emergence of new language models, specifically BERT, there is hope for improvements. The better word embedding capabilities of BERT lend to this study, aiding in the capturing of the nuanced semantic pool and the contextual intricacies of the language under study, the Kurdish language, thus setting a new benchmark for sentiment analysis in low-resource languages.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16804
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KuBERT: Central Kurdish BERT Model and Its Application for Sentiment Analysis
Awlla, Kozhin muhealddin
Veisi, Hadi
Abdullah, Abdulhady Abas
Computation and Language
Artificial Intelligence
This paper enhances the study of sentiment analysis for the Central Kurdish language by integrating the Bidirectional Encoder Representations from Transformers (BERT) into Natural Language Processing techniques. Kurdish is a low-resourced language, having a high level of linguistic diversity with minimal computational resources, making sentiment analysis somewhat challenging. Earlier, this was done using a traditional word embedding model, such as Word2Vec, but with the emergence of new language models, specifically BERT, there is hope for improvements. The better word embedding capabilities of BERT lend to this study, aiding in the capturing of the nuanced semantic pool and the contextual intricacies of the language under study, the Kurdish language, thus setting a new benchmark for sentiment analysis in low-resource languages.
title KuBERT: Central Kurdish BERT Model and Its Application for Sentiment Analysis
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.16804