AlbNews: A Corpus of Headlines for Topic Modeling in Albanian

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Çano, Erion, Lamaj, Dario
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913224886059008
author Çano, Erion
Lamaj, Dario
author_facet Çano, Erion
Lamaj, Dario
contents The scarcity of available text corpora for low-resource languages like Albanian is a serious hurdle for research in natural language processing tasks. This paper introduces AlbNews, a collection of 600 topically labeled news headlines and 2600 unlabeled ones in Albanian. The data can be freely used for conducting topic modeling research. We report the initial classification scores of some traditional machine learning classifiers trained with the AlbNews samples. These results show that basic models outrun the ensemble learning ones and can serve as a baseline for future experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04028
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AlbNews: A Corpus of Headlines for Topic Modeling in Albanian
Çano, Erion
Lamaj, Dario
Computation and Language
Artificial Intelligence
Machine Learning
The scarcity of available text corpora for low-resource languages like Albanian is a serious hurdle for research in natural language processing tasks. This paper introduces AlbNews, a collection of 600 topically labeled news headlines and 2600 unlabeled ones in Albanian. The data can be freely used for conducting topic modeling research. We report the initial classification scores of some traditional machine learning classifiers trained with the AlbNews samples. These results show that basic models outrun the ensemble learning ones and can serve as a baseline for future experiments.
title AlbNews: A Corpus of Headlines for Topic Modeling in Albanian
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2402.04028