Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912604304179200 |
|---|---|
| author | Zhang, Tuo Sun, Yuechun Liu, Ruiliang |
| author_facet | Zhang, Tuo Sun, Yuechun Liu, Ruiliang |
| contents | In this work, we present a retrieval-augmented generation (RAG)-based system for provenance analysis of archaeological artifacts, designed to support expert reasoning by integrating multimodal retrieval and large vision-language models (VLMs). The system constructs a dual-modal knowledge base from reference texts and images, enabling raw visual, edge-enhanced, and semantic retrieval to identify stylistically similar objects. Retrieved candidates are synthesized by the VLM to generate structured inferences, including chronological, geographical, and cultural attributions, alongside interpretive justifications. We evaluate the system on a set of Eastern Eurasian Bronze Age artifacts from the British Museum. Expert evaluation demonstrates that the system produces meaningful and interpretable outputs, offering scholars concrete starting points for analysis and significantly alleviating the cognitive burden of navigating vast comparative corpora. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_20769 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems Zhang, Tuo Sun, Yuechun Liu, Ruiliang Information Retrieval Artificial Intelligence Computer Vision and Pattern Recognition In this work, we present a retrieval-augmented generation (RAG)-based system for provenance analysis of archaeological artifacts, designed to support expert reasoning by integrating multimodal retrieval and large vision-language models (VLMs). The system constructs a dual-modal knowledge base from reference texts and images, enabling raw visual, edge-enhanced, and semantic retrieval to identify stylistically similar objects. Retrieved candidates are synthesized by the VLM to generate structured inferences, including chronological, geographical, and cultural attributions, alongside interpretive justifications. We evaluate the system on a set of Eastern Eurasian Bronze Age artifacts from the British Museum. Expert evaluation demonstrates that the system produces meaningful and interpretable outputs, offering scholars concrete starting points for analysis and significantly alleviating the cognitive burden of navigating vast comparative corpora. |
| title | Provenance Analysis of Archaeological Artifacts via Multimodal RAG Systems |
| topic | Information Retrieval Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2509.20769 |