Metadata Enrichment of Long Text Documents using Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lamba, Manika, Peng, You, Nikolov, Sophie, Layne-Worthey, Glen, Downie, J. Stephen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913912981553152
author Lamba, Manika
Peng, You
Nikolov, Sophie
Layne-Worthey, Glen
Downie, J. Stephen
author_facet Lamba, Manika
Peng, You
Nikolov, Sophie
Layne-Worthey, Glen
Downie, J. Stephen
contents In this project, we semantically enriched and enhanced the metadata of long text documents, theses and dissertations, retrieved from the HathiTrust Digital Library in English published from 1920 to 2020 through a combination of manual efforts and large language models. This dataset provides a valuable resource for advancing research in areas such as computational social science, digital humanities, and information science. Our paper shows that enriching metadata using LLMs is particularly beneficial for digital repositories by introducing additional metadata access points that may not have originally been foreseen to accommodate various content types. This approach is particularly effective for repositories that have significant missing data in their existing metadata fields, enhancing search results and improving the accessibility of the digital repository.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20918
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Metadata Enrichment of Long Text Documents using Large Language Models
Lamba, Manika
Peng, You
Nikolov, Sophie
Layne-Worthey, Glen
Downie, J. Stephen
Digital Libraries
Emerging Technologies
Information Retrieval
In this project, we semantically enriched and enhanced the metadata of long text documents, theses and dissertations, retrieved from the HathiTrust Digital Library in English published from 1920 to 2020 through a combination of manual efforts and large language models. This dataset provides a valuable resource for advancing research in areas such as computational social science, digital humanities, and information science. Our paper shows that enriching metadata using LLMs is particularly beneficial for digital repositories by introducing additional metadata access points that may not have originally been foreseen to accommodate various content types. This approach is particularly effective for repositories that have significant missing data in their existing metadata fields, enhancing search results and improving the accessibility of the digital repository.
title Metadata Enrichment of Long Text Documents using Large Language Models
topic Digital Libraries
Emerging Technologies
Information Retrieval
url https://arxiv.org/abs/2506.20918