Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Zhiyin, D'Souza, Jennifer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910824058060800
author Tan, Zhiyin
D'Souza, Jennifer
author_facet Tan, Zhiyin
D'Souza, Jennifer
contents This study presents a framework for automated evaluation of dynamically evolving topic taxonomies in scientific literature using Large Language Models (LLMs). In digital library systems, topic modeling plays a crucial role in efficiently organizing and retrieving scholarly content, guiding researchers through complex knowledge landscapes. As research domains proliferate and shift, traditional human centric and static evaluation methods struggle to maintain relevance. The proposed approach harnesses LLMs to measure key quality dimensions, such as coherence, repetitiveness, diversity, and topic-document alignment, without heavy reliance on expert annotators or narrow statistical metrics. Tailored prompts guide LLM assessments, ensuring consistent and interpretable evaluations across various datasets and modeling techniques. Experiments on benchmark corpora demonstrate the method's robustness, scalability, and adaptability, underscoring its value as a more holistic and dynamic alternative to conventional evaluation strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07352
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation
Tan, Zhiyin
D'Souza, Jennifer
Computation and Language
Artificial Intelligence
Digital Libraries
This study presents a framework for automated evaluation of dynamically evolving topic taxonomies in scientific literature using Large Language Models (LLMs). In digital library systems, topic modeling plays a crucial role in efficiently organizing and retrieving scholarly content, guiding researchers through complex knowledge landscapes. As research domains proliferate and shift, traditional human centric and static evaluation methods struggle to maintain relevance. The proposed approach harnesses LLMs to measure key quality dimensions, such as coherence, repetitiveness, diversity, and topic-document alignment, without heavy reliance on expert annotators or narrow statistical metrics. Tailored prompts guide LLM assessments, ensuring consistent and interpretable evaluations across various datasets and modeling techniques. Experiments on benchmark corpora demonstrate the method's robustness, scalability, and adaptability, underscoring its value as a more holistic and dynamic alternative to conventional evaluation strategies.
title Bridging the Evaluation Gap: Leveraging Large Language Models for Topic Model Evaluation
topic Computation and Language
Artificial Intelligence
Digital Libraries
url https://arxiv.org/abs/2502.07352