Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Medvecki, Darija, Bašaragin, Bojana, Ljajić, Adela, Milošević, Nikola
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916115646513152
author Medvecki, Darija
Bašaragin, Bojana
Ljajić, Adela
Milošević, Nikola
author_facet Medvecki, Darija
Bašaragin, Bojana
Ljajić, Adela
Milošević, Nikola
contents This paper presents the results of the first application of BERTopic, a state-of-the-art topic modeling technique, to short text written in a morphologi-cally rich language. We applied BERTopic with three multilingual embed-ding models on two levels of text preprocessing (partial and full) to evalu-ate its performance on partially preprocessed short text in Serbian. We also compared it to LDA and NMF on fully preprocessed text. The experiments were conducted on a dataset of tweets expressing hesitancy toward COVID-19 vaccination. Our results show that with adequate parameter setting, BERTopic can yield informative topics even when applied to partially pre-processed short text. When the same parameters are applied in both prepro-cessing scenarios, the performance drop on partially preprocessed text is minimal. Compared to LDA and NMF, judging by the keywords, BERTopic offers more informative topics and gives novel insights when the number of topics is not limited. The findings of this paper can be significant for re-searchers working with other morphologically rich low-resource languages and short text.
format Preprint
id arxiv_https___arxiv_org_abs_2402_03067
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian
Medvecki, Darija
Bašaragin, Bojana
Ljajić, Adela
Milošević, Nikola
Computation and Language
Artificial Intelligence
This paper presents the results of the first application of BERTopic, a state-of-the-art topic modeling technique, to short text written in a morphologi-cally rich language. We applied BERTopic with three multilingual embed-ding models on two levels of text preprocessing (partial and full) to evalu-ate its performance on partially preprocessed short text in Serbian. We also compared it to LDA and NMF on fully preprocessed text. The experiments were conducted on a dataset of tweets expressing hesitancy toward COVID-19 vaccination. Our results show that with adequate parameter setting, BERTopic can yield informative topics even when applied to partially pre-processed short text. When the same parameters are applied in both prepro-cessing scenarios, the performance drop on partially preprocessed text is minimal. Compared to LDA and NMF, judging by the keywords, BERTopic offers more informative topics and gives novel insights when the number of topics is not limited. The findings of this paper can be significant for re-searchers working with other morphologically rich low-resource languages and short text.
title Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.03067