Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Jinyu, Song, Xiaoying, Zhang, Diana, Thomale, Jason, He, Daqing, Hong, Lingzi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2507.22913
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913967084929024
author Liu, Jinyu
Song, Xiaoying
Zhang, Diana
Thomale, Jason
He, Daqing
Hong, Lingzi
author_facet Liu, Jinyu
Song, Xiaoying
Zhang, Diana
Thomale, Jason
He, Daqing
Hong, Lingzi
contents Providing subject access to information resources is an essential function of any library management system. Large language models (LLMs) have been widely used in classification and summarization tasks, but their capability to perform subject analysis is underexplored. Multi-label classification with traditional machine learning (ML) models has been used for subject analysis but struggles with unseen cases. LLMs offer an alternative but often over-generate and hallucinate. Therefore, we propose a hybrid framework that integrates embedding-based ML models with LLMs. This approach uses ML models to (1) predict the optimal number of LCSH labels to guide LLM predictions and (2) post-edit the predicted terms with actual LCSH terms to mitigate hallucinations. We experimented with LLMs and the hybrid framework to predict the subject terms of books using the Library of Congress Subject Headings (LCSH). Experiment results show that providing initial predictions to guide LLM generations and imposing post-edits result in more controlled and vocabulary-aligned outputs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models
Liu, Jinyu
Song, Xiaoying
Zhang, Diana
Thomale, Jason
He, Daqing
Hong, Lingzi
Computation and Language
Artificial Intelligence
Providing subject access to information resources is an essential function of any library management system. Large language models (LLMs) have been widely used in classification and summarization tasks, but their capability to perform subject analysis is underexplored. Multi-label classification with traditional machine learning (ML) models has been used for subject analysis but struggles with unseen cases. LLMs offer an alternative but often over-generate and hallucinate. Therefore, we propose a hybrid framework that integrates embedding-based ML models with LLMs. This approach uses ML models to (1) predict the optimal number of LCSH labels to guide LLM predictions and (2) post-edit the predicted terms with actual LCSH terms to mitigate hallucinations. We experimented with LLMs and the hybrid framework to predict the subject terms of books using the Library of Congress Subject Headings (LCSH). Experiment results show that providing initial predictions to guide LLM generations and imposing post-edits result in more controlled and vocabulary-aligned outputs.
title A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.22913