LLM as Attention-Informed NTM and Topic Modeling as long-input Generation: Interpretability and long-Context Capability

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Xuan, Yang, Zhongliang, Li, Haolun, Chu, Beilin, Tian, Rui, Li, Yu, Tan, Shaolin, Zhou, Linna
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910126962638848
author Xu, Xuan
Yang, Zhongliang
Li, Haolun
Chu, Beilin
Tian, Rui
Li, Yu
Tan, Shaolin
Zhou, Linna
author_facet Xu, Xuan
Yang, Zhongliang
Li, Haolun
Chu, Beilin
Tian, Rui
Li, Yu
Tan, Shaolin
Zhou, Linna
contents Topic modeling aims to produce interpretable topic representations and topic--document correspondences from corpora, but classical neural topic models (NTMs) remain constrained by limited representation assumptions and semantic abstraction ability. We study LLM-based topic modeling from both white-box and black-box perspectives. For white-box LLMs, we propose an attention-informed framework that recovers interpretable structures analogous to those in NTMs, including document-topic and topic-word distributions. This validates the view that LLM can serve as an attention-informed NTM. For black-box LLMs, we reformulate topic modeling as a structured long-input task and introduce a post-generation signal compensation method based on diversified topic cues and hybrid retrieval. Experiments show that recovered attention structures support effective topic assignment and keyword extraction, while black-box long-context LLMs achieve competitive or stronger performance than other baselines. These findings suggest a connection between LLMs and NTMs and highlight the promise of long-context LLMs for topic modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03174
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM as Attention-Informed NTM and Topic Modeling as long-input Generation: Interpretability and long-Context Capability
Xu, Xuan
Yang, Zhongliang
Li, Haolun
Chu, Beilin
Tian, Rui
Li, Yu
Tan, Shaolin
Zhou, Linna
Computation and Language
Artificial Intelligence
Topic modeling aims to produce interpretable topic representations and topic--document correspondences from corpora, but classical neural topic models (NTMs) remain constrained by limited representation assumptions and semantic abstraction ability. We study LLM-based topic modeling from both white-box and black-box perspectives. For white-box LLMs, we propose an attention-informed framework that recovers interpretable structures analogous to those in NTMs, including document-topic and topic-word distributions. This validates the view that LLM can serve as an attention-informed NTM. For black-box LLMs, we reformulate topic modeling as a structured long-input task and introduce a post-generation signal compensation method based on diversified topic cues and hybrid retrieval. Experiments show that recovered attention structures support effective topic assignment and keyword extraction, while black-box long-context LLMs achieve competitive or stronger performance than other baselines. These findings suggest a connection between LLMs and NTMs and highlight the promise of long-context LLMs for topic modeling.
title LLM as Attention-Informed NTM and Topic Modeling as long-input Generation: Interpretability and long-Context Capability
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.03174