VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Jinchao, Cheng, Tengfei, Wu, Biao, Zhang, Zeyu, Huang, Shiya, Bishop, Judith, Shepherd, Gillian, Fang, Meng, Chen, Ling, Zhao, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911394929049600
author Ge, Jinchao
Cheng, Tengfei
Wu, Biao
Zhang, Zeyu
Huang, Shiya
Bishop, Judith
Shepherd, Gillian
Fang, Meng
Chen, Ling
Zhao, Yang
author_facet Ge, Jinchao
Cheng, Tengfei
Wu, Biao
Zhang, Zeyu
Huang, Shiya
Bishop, Judith
Shepherd, Gillian
Fang, Meng
Chen, Ling
Zhao, Yang
contents Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific data. We introduce VaseVQA, a benchmark of 31,773 images and 67,614 question-answer pairs across seven expert-defined categories, enabling systematic evaluation of expert-level cultural heritage understanding. Using this dataset, we explore effective training strategies for domain-specific reasoning. While supervised fine-tuning improves adaptation to domain knowledge, it struggles with deeper reasoning tasks. We propose VaseVL, which augments SFT with reinforcement learning using verifiable rewards. Experiments show that VaseVL consistently outperforms supervised baselines, especially on reasoning-intensive questions, highlighting the value of targeted reinforcement learning for cultural heritage visual question answering. Our code and dataset will be released at https://github.com/AIGeeksGroup/VaseVQA.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17191
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
Ge, Jinchao
Cheng, Tengfei
Wu, Biao
Zhang, Zeyu
Huang, Shiya
Bishop, Judith
Shepherd, Gillian
Fang, Meng
Chen, Ling
Zhao, Yang
Computer Vision and Pattern Recognition
Computation and Language
Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific data. We introduce VaseVQA, a benchmark of 31,773 images and 67,614 question-answer pairs across seven expert-defined categories, enabling systematic evaluation of expert-level cultural heritage understanding. Using this dataset, we explore effective training strategies for domain-specific reasoning. While supervised fine-tuning improves adaptation to domain knowledge, it struggles with deeper reasoning tasks. We propose VaseVL, which augments SFT with reinforcement learning using verifiable rewards. Experiments show that VaseVL consistently outperforms supervised baselines, especially on reasoning-intensive questions, highlighting the value of targeted reinforcement learning for cultural heritage visual question answering. Our code and dataset will be released at https://github.com/AIGeeksGroup/VaseVQA.
title VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2509.17191