VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911394929049600 |
|---|---|
| author | Ge, Jinchao Cheng, Tengfei Wu, Biao Zhang, Zeyu Huang, Shiya Bishop, Judith Shepherd, Gillian Fang, Meng Chen, Ling Zhao, Yang |
| author_facet | Ge, Jinchao Cheng, Tengfei Wu, Biao Zhang, Zeyu Huang, Shiya Bishop, Judith Shepherd, Gillian Fang, Meng Chen, Ling Zhao, Yang |
| contents | Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific data. We introduce VaseVQA, a benchmark of 31,773 images and 67,614 question-answer pairs across seven expert-defined categories, enabling systematic evaluation of expert-level cultural heritage understanding. Using this dataset, we explore effective training strategies for domain-specific reasoning. While supervised fine-tuning improves adaptation to domain knowledge, it struggles with deeper reasoning tasks. We propose VaseVL, which augments SFT with reinforcement learning using verifiable rewards. Experiments show that VaseVL consistently outperforms supervised baselines, especially on reasoning-intensive questions, highlighting the value of targeted reinforcement learning for cultural heritage visual question answering. Our code and dataset will be released at https://github.com/AIGeeksGroup/VaseVQA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_17191 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery Ge, Jinchao Cheng, Tengfei Wu, Biao Zhang, Zeyu Huang, Shiya Bishop, Judith Shepherd, Gillian Fang, Meng Chen, Ling Zhao, Yang Computer Vision and Pattern Recognition Computation and Language Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific data. We introduce VaseVQA, a benchmark of 31,773 images and 67,614 question-answer pairs across seven expert-defined categories, enabling systematic evaluation of expert-level cultural heritage understanding. Using this dataset, we explore effective training strategies for domain-specific reasoning. While supervised fine-tuning improves adaptation to domain knowledge, it struggles with deeper reasoning tasks. We propose VaseVL, which augments SFT with reinforcement learning using verifiable rewards. Experiments show that VaseVL consistently outperforms supervised baselines, especially on reasoning-intensive questions, highlighting the value of targeted reinforcement learning for cultural heritage visual question answering. Our code and dataset will be released at https://github.com/AIGeeksGroup/VaseVQA. |
| title | VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery |
| topic | Computer Vision and Pattern Recognition Computation and Language |
| url | https://arxiv.org/abs/2509.17191 |