Leveraging Language Models for Interpretable Analysis of Narratives in a Large Corpus

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bai, Eric A., Zhou, Minling, Henao, Ricardo, Schwing, Kyle M., Carin, Lawrence
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909920036651008
author Bai, Eric A.
Zhou, Minling
Henao, Ricardo
Schwing, Kyle M.
Carin, Lawrence
author_facet Bai, Eric A.
Zhou, Minling
Henao, Ricardo
Schwing, Kyle M.
Carin, Lawrence
contents Narratives drive human behavior and lay at the core of geopolitics, but have eluded quantification that would permit measurement of their overlap and evolution. We present an interpretable model that integrates an established bag-of-words (BoW) topical representation and a novel LLM-based question answering (Q&A) narrative model, which share a latent Reproducing Kernel Hilbert Space representation, to quantify written documents. Our approach mitigates the cost, interpretability, and generalization challenges of using a LLM to analyze large corpora without full inference. We derive efficient functional gradient descent updates that are interpretable and structurally analogous to the self-attention mechanism in Transformers. We further introduce an in-context Q&A extrapolation method inspired by Transformer architectures, enabling accurate prediction of Q&A outcomes for unqueried documents.
format Preprint
id arxiv_https___arxiv_org_abs_2511_18599
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Leveraging Language Models for Interpretable Analysis of Narratives in a Large Corpus
Bai, Eric A.
Zhou, Minling
Henao, Ricardo
Schwing, Kyle M.
Carin, Lawrence
Signal Processing
Narratives drive human behavior and lay at the core of geopolitics, but have eluded quantification that would permit measurement of their overlap and evolution. We present an interpretable model that integrates an established bag-of-words (BoW) topical representation and a novel LLM-based question answering (Q&A) narrative model, which share a latent Reproducing Kernel Hilbert Space representation, to quantify written documents. Our approach mitigates the cost, interpretability, and generalization challenges of using a LLM to analyze large corpora without full inference. We derive efficient functional gradient descent updates that are interpretable and structurally analogous to the self-attention mechanism in Transformers. We further introduce an in-context Q&A extrapolation method inspired by Transformer architectures, enabling accurate prediction of Q&A outcomes for unqueried documents.
title Leveraging Language Models for Interpretable Analysis of Narratives in a Large Corpus
topic Signal Processing
url https://arxiv.org/abs/2511.18599