ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Abele, Line, Anders, Gerrit, Aydın, Tolgahan, Buder, Jürgen, Fischer, Helen, Kimmel, Dominik, Huff, Markus
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915381738733568
author Abele, Line
Anders, Gerrit
Aydın, Tolgahan
Buder, Jürgen
Fischer, Helen
Kimmel, Dominik
Huff, Markus
author_facet Abele, Line
Anders, Gerrit
Aydın, Tolgahan
Buder, Jürgen
Fischer, Helen
Kimmel, Dominik
Huff, Markus
contents The accelerating growth of photographic collections has outpaced manual cataloguing, motivating the use of vision language models (VLMs) to automate metadata generation. This study examines whether Al-generated catalogue descriptions can approximate human-written quality and how generative Al might integrate into cataloguing workflows in archival and museum collections. A VLM (InternVL2) generated catalogue descriptions for photographic prints on labelled cardboard mounts with archaeological content, evaluated by archive and archaeology experts and non-experts in a human-centered, experimental framework. Participants classified descriptions as AI-generated or expert-written, rated quality, and reported willingness to use and trust in AI tools. Classification performance was above chance level, with both groups underestimating their ability to detect Al-generated descriptions. OCR errors and hallucinations limited perceived quality, yet descriptions rated higher in accuracy and usefulness were harder to classify, suggesting that human review is necessary to ensure the accuracy and quality of catalogue descriptions generated by the out-of-the-box model, particularly in specialized domains like archaeological cataloguing. Experts showed lower willingness to adopt AI tools, emphasizing concerns on preservation responsibility over technical performance. These findings advocate for a collaborative approach where AI supports draft generation but remains subordinate to human verification, ensuring alignment with curatorial values (e.g., provenance, transparency). The successful integration of this approach depends not only on technical advancements, such as domain-specific fine-tuning, but even more on establishing trust among professionals, which could both be fostered through a transparent and explainable AI pipeline.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07551
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing
Abele, Line
Anders, Gerrit
Aydın, Tolgahan
Buder, Jürgen
Fischer, Helen
Kimmel, Dominik
Huff, Markus
Human-Computer Interaction
Artificial Intelligence
Digital Libraries
The accelerating growth of photographic collections has outpaced manual cataloguing, motivating the use of vision language models (VLMs) to automate metadata generation. This study examines whether Al-generated catalogue descriptions can approximate human-written quality and how generative Al might integrate into cataloguing workflows in archival and museum collections. A VLM (InternVL2) generated catalogue descriptions for photographic prints on labelled cardboard mounts with archaeological content, evaluated by archive and archaeology experts and non-experts in a human-centered, experimental framework. Participants classified descriptions as AI-generated or expert-written, rated quality, and reported willingness to use and trust in AI tools. Classification performance was above chance level, with both groups underestimating their ability to detect Al-generated descriptions. OCR errors and hallucinations limited perceived quality, yet descriptions rated higher in accuracy and usefulness were harder to classify, suggesting that human review is necessary to ensure the accuracy and quality of catalogue descriptions generated by the out-of-the-box model, particularly in specialized domains like archaeological cataloguing. Experts showed lower willingness to adopt AI tools, emphasizing concerns on preservation responsibility over technical performance. These findings advocate for a collaborative approach where AI supports draft generation but remains subordinate to human verification, ensuring alignment with curatorial values (e.g., provenance, transparency). The successful integration of this approach depends not only on technical advancements, such as domain-specific fine-tuning, but even more on establishing trust among professionals, which could both be fostered through a transparent and explainable AI pipeline.
title ArchiveGPT: A human-centered evaluation of using a vision language model for image cataloguing
topic Human-Computer Interaction
Artificial Intelligence
Digital Libraries
url https://arxiv.org/abs/2507.07551