LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Jian, Guo, Junyi, Yuan, Junyi, Lu, Huanda, Zhou, Yanlin, Wu, Fangyu, Wang, Qiufeng, Lu, Dongming
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908639103549440
author Zhang, Jian
Guo, Junyi
Yuan, Junyi
Lu, Huanda
Zhou, Yanlin
Wu, Fangyu
Wang, Qiufeng
Lu, Dongming
author_facet Zhang, Jian
Guo, Junyi
Yuan, Junyi
Lu, Huanda
Zhou, Yanlin
Wu, Fangyu
Wang, Qiufeng
Lu, Dongming
contents Cross-modal retrieval is essential for interpreting cultural heritage data, but its effectiveness is often limited by incomplete or inconsistent textual descriptions, caused by historical data loss and the high cost of expert annotation. While large language models (LLMs) offer a promising solution by enriching textual descriptions, their outputs frequently suffer from hallucinations or miss visually grounded details. To address these challenges, we propose $C^3$, a data augmentation framework that enhances cross-modal retrieval performance by improving the completeness and consistency of LLM-generated descriptions. $C^3$ introduces a completeness evaluation module to assess semantic coverage using both visual cues and language-model outputs. Furthermore, to mitigate factual inconsistencies, we formulate a Markov Decision Process to supervise Chain-of-Thought reasoning, guiding consistency evaluation through adaptive query control. Experiments on the cultural heritage datasets CulTi and TimeTravel, as well as on general benchmarks MSCOCO and Flickr30K, demonstrate that $C^3$ achieves state-of-the-art performance in both fine-tuned and zero-shot settings.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06268
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval
Zhang, Jian
Guo, Junyi
Yuan, Junyi
Lu, Huanda
Zhou, Yanlin
Wu, Fangyu
Wang, Qiufeng
Lu, Dongming
Computer Vision and Pattern Recognition
Computers and Society
Cross-modal retrieval is essential for interpreting cultural heritage data, but its effectiveness is often limited by incomplete or inconsistent textual descriptions, caused by historical data loss and the high cost of expert annotation. While large language models (LLMs) offer a promising solution by enriching textual descriptions, their outputs frequently suffer from hallucinations or miss visually grounded details. To address these challenges, we propose $C^3$, a data augmentation framework that enhances cross-modal retrieval performance by improving the completeness and consistency of LLM-generated descriptions. $C^3$ introduces a completeness evaluation module to assess semantic coverage using both visual cues and language-model outputs. Furthermore, to mitigate factual inconsistencies, we formulate a Markov Decision Process to supervise Chain-of-Thought reasoning, guiding consistency evaluation through adaptive query control. Experiments on the cultural heritage datasets CulTi and TimeTravel, as well as on general benchmarks MSCOCO and Flickr30K, demonstrate that $C^3$ achieves state-of-the-art performance in both fine-tuned and zero-shot settings.
title LLM-Driven Completeness and Consistency Evaluation for Cultural Heritage Data Augmentation in Cross-Modal Retrieval
topic Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2511.06268