VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Nonghai, Zhang, Zeyu, Wang, Jiazi, Zhao, Yang, Tang, Hao
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914087342964736
author Zhang, Nonghai
Zhang, Zeyu
Wang, Jiazi
Zhao, Yang
Tang, Hao
author_facet Zhang, Nonghai
Zhang, Zeyu
Wang, Jiazi
Zhao, Yang
Tang, Hao
contents Vision-Language Models (VLMs) have achieved significant progress in multimodal understanding tasks, demonstrating strong capabilities particularly in general tasks such as image captioning and visual reasoning. However, when dealing with specialized cultural heritage domains like 3D vase artifacts, existing models face severe data scarcity issues and insufficient domain knowledge limitations. Due to the lack of targeted training data, current VLMs struggle to effectively handle such culturally significant specialized tasks. To address these challenges, we propose the VaseVQA-3D dataset, which serves as the first 3D visual question answering dataset for ancient Greek pottery analysis, collecting 664 ancient Greek vase 3D models with corresponding question-answer data and establishing a complete data construction pipeline. We further develop the VaseVLM model, enhancing model performance in vase artifact analysis through domain-adaptive training. Experimental results validate the effectiveness of our approach, where we improve by 12.8% on R@1 metrics and by 6.6% on lexical similarity compared with previous state-of-the-art on the VaseVQA-3D dataset, significantly improving the recognition and understanding of 3D vase artifacts, providing new technical pathways for digital heritage preservation research. Code: https://github.com/AIGeeksGroup/VaseVQA-3D. Website: https://aigeeksgroup.github.io/VaseVQA-3D.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04479
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
Zhang, Nonghai
Zhang, Zeyu
Wang, Jiazi
Zhao, Yang
Tang, Hao
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) have achieved significant progress in multimodal understanding tasks, demonstrating strong capabilities particularly in general tasks such as image captioning and visual reasoning. However, when dealing with specialized cultural heritage domains like 3D vase artifacts, existing models face severe data scarcity issues and insufficient domain knowledge limitations. Due to the lack of targeted training data, current VLMs struggle to effectively handle such culturally significant specialized tasks. To address these challenges, we propose the VaseVQA-3D dataset, which serves as the first 3D visual question answering dataset for ancient Greek pottery analysis, collecting 664 ancient Greek vase 3D models with corresponding question-answer data and establishing a complete data construction pipeline. We further develop the VaseVLM model, enhancing model performance in vase artifact analysis through domain-adaptive training. Experimental results validate the effectiveness of our approach, where we improve by 12.8% on R@1 metrics and by 6.6% on lexical similarity compared with previous state-of-the-art on the VaseVQA-3D dataset, significantly improving the recognition and understanding of 3D vase artifacts, providing new technical pathways for digital heritage preservation research. Code: https://github.com/AIGeeksGroup/VaseVQA-3D. Website: https://aigeeksgroup.github.io/VaseVQA-3D.
title VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.04479