Semantic-Augmented Latent Topic Modeling with LLM-in-the-Loop

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hong, Mengze, Zhang, Chen Jason, Jiang, Di
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915383359832064
author Hong, Mengze
Zhang, Chen Jason
Jiang, Di
author_facet Hong, Mengze
Zhang, Chen Jason
Jiang, Di
contents Latent Dirichlet Allocation (LDA) is a prominent generative probabilistic model used for uncovering abstract topics within document collections. In this paper, we explore the effectiveness of augmenting topic models with Large Language Models (LLMs) through integration into two key phases: Initialization and Post-Correction. Since the LDA is highly dependent on the quality of its initialization, we conduct extensive experiments on the LLM-guided topic clustering for initializing the Gibbs sampling algorithm. Interestingly, the experimental results reveal that while the proposed initialization strategy improves the early iterations of LDA, it has no effect on the convergence and yields the worst performance compared to the baselines. The LLM-enabled post-correction, on the other hand, achieved a promising improvement of 5.86% in the coherence evaluation. These results highlight the practical benefits of the LLM-in-the-loop approach and challenge the belief that LLMs are always the superior text mining alternative.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08498
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semantic-Augmented Latent Topic Modeling with LLM-in-the-Loop
Hong, Mengze
Zhang, Chen Jason
Jiang, Di
Computation and Language
Latent Dirichlet Allocation (LDA) is a prominent generative probabilistic model used for uncovering abstract topics within document collections. In this paper, we explore the effectiveness of augmenting topic models with Large Language Models (LLMs) through integration into two key phases: Initialization and Post-Correction. Since the LDA is highly dependent on the quality of its initialization, we conduct extensive experiments on the LLM-guided topic clustering for initializing the Gibbs sampling algorithm. Interestingly, the experimental results reveal that while the proposed initialization strategy improves the early iterations of LDA, it has no effect on the convergence and yields the worst performance compared to the baselines. The LLM-enabled post-correction, on the other hand, achieved a promising improvement of 5.86% in the coherence evaluation. These results highlight the practical benefits of the LLM-in-the-loop approach and challenge the belief that LLMs are always the superior text mining alternative.
title Semantic-Augmented Latent Topic Modeling with LLM-in-the-Loop
topic Computation and Language
url https://arxiv.org/abs/2507.08498