BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mutsaddi, Atharva, Jamkhande, Anvi, Thakre, Aryan, Haribhakta, Yashodhara
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910775835099136
author Mutsaddi, Atharva
Jamkhande, Anvi
Thakre, Aryan
Haribhakta, Yashodhara
author_facet Mutsaddi, Atharva
Jamkhande, Anvi
Thakre, Aryan
Haribhakta, Yashodhara
contents As short text data in native languages like Hindi increasingly appear in modern media, robust methods for topic modeling on such data have gained importance. This study investigates the performance of BERTopic in modeling Hindi short texts, an area that has been under-explored in existing research. Using contextual embeddings, BERTopic can capture semantic relationships in data, making it potentially more effective than traditional models, especially for short and diverse texts. We evaluate BERTopic using 6 different document embedding models and compare its performance against 8 established topic modeling techniques, such as Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorization (NMF), Latent Semantic Indexing (LSI), Additive Regularization of Topic Models (ARTM), Probabilistic Latent Semantic Analysis (PLSA), Embedded Topic Model (ETM), Combined Topic Model (CTM), and Top2Vec. The models are assessed using coherence scores across a range of topic counts. Our results reveal that BERTopic consistently outperforms other models in capturing coherent topics from short Hindi texts.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03843
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study
Mutsaddi, Atharva
Jamkhande, Anvi
Thakre, Aryan
Haribhakta, Yashodhara
Information Retrieval
Computation and Language
Machine Learning
I.2.7; H.3.3; H.2.8
As short text data in native languages like Hindi increasingly appear in modern media, robust methods for topic modeling on such data have gained importance. This study investigates the performance of BERTopic in modeling Hindi short texts, an area that has been under-explored in existing research. Using contextual embeddings, BERTopic can capture semantic relationships in data, making it potentially more effective than traditional models, especially for short and diverse texts. We evaluate BERTopic using 6 different document embedding models and compare its performance against 8 established topic modeling techniques, such as Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorization (NMF), Latent Semantic Indexing (LSI), Additive Regularization of Topic Models (ARTM), Probabilistic Latent Semantic Analysis (PLSA), Embedded Topic Model (ETM), Combined Topic Model (CTM), and Top2Vec. The models are assessed using coherence scores across a range of topic counts. Our results reveal that BERTopic consistently outperforms other models in capturing coherent topics from short Hindi texts.
title BERTopic for Topic Modeling of Hindi Short Texts: A Comparative Study
topic Information Retrieval
Computation and Language
Machine Learning
I.2.7; H.3.3; H.2.8
url https://arxiv.org/abs/2501.03843