VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
Fuente:
arXiv
Saved in:
| Main Authors: | Ge, Jinchao, Cheng, Tengfei, Wu, Biao, Zhang, Zeyu, Huang, Shiya, Bishop, Judith, Shepherd, Gillian, Fang, Meng, Chen, Ling, Zhao, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025)
by: Zhang, Nonghai, et al.
Published: (2025)
Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
by: Yan, Junfeng, et al.
Published: (2025)
by: Yan, Junfeng, et al.
Published: (2025)
PresentAgent: Multimodal Agent for Presentation Video Generation
by: Shi, Jingwei, et al.
Published: (2025)
by: Shi, Jingwei, et al.
Published: (2025)
Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models
by: Mavromatis, Spyridon, et al.
Published: (2026)
by: Mavromatis, Spyridon, et al.
Published: (2026)
Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
by: Li, Yangning, et al.
Published: (2024)
by: Li, Yangning, et al.
Published: (2024)
Enhancing Human Pose Estimation in Ancient Vase Paintings via Perceptually-grounded Style Transfer Learning
by: Madhu, Prathmesh, et al.
Published: (2020)
by: Madhu, Prathmesh, et al.
Published: (2020)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
by: Cheng, Xianfu, et al.
Published: (2025)
by: Cheng, Xianfu, et al.
Published: (2025)
A State-of-the-Art Morphosyntactic Parser and Lemmatizer for Ancient Greek
by: Celano, Giuseppe G. A.
Published: (2024)
by: Celano, Giuseppe G. A.
Published: (2024)
EpiAgent: An Agent-Centric System for Ancient Inscription Restoration
by: Zhu, Shipeng, et al.
Published: (2026)
by: Zhu, Shipeng, et al.
Published: (2026)
Contrastive Learning for Character Detection in Ancient Greek Papyri
by: Nakka, Vedasri, et al.
Published: (2024)
by: Nakka, Vedasri, et al.
Published: (2024)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
by: Baek, Jeonghun, et al.
Published: (2025)
by: Baek, Jeonghun, et al.
Published: (2025)
Artificial Intelligence in Education: Ethical Considerations and Insights from Ancient Greek Philosophy
by: Karpouzis, Kostas
Published: (2024)
by: Karpouzis, Kostas
Published: (2024)
Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling
by: Zou, Hongjian, et al.
Published: (2026)
by: Zou, Hongjian, et al.
Published: (2026)
MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents
by: Yang, Wanqi, et al.
Published: (2025)
by: Yang, Wanqi, et al.
Published: (2025)
MMA: Multimodal Memory Agent
by: Lu, Yihao, et al.
Published: (2026)
by: Lu, Yihao, et al.
Published: (2026)
MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
KNVQA: A Benchmark for evaluation knowledge-based VQA
by: Cheng, Sirui, et al.
Published: (2023)
by: Cheng, Sirui, et al.
Published: (2023)
UVA 3D Greek Vases - aryballos
by: UVA3D
Published: (2021)
by: UVA3D
Published: (2021)
An Embedded Diachronic Sense Change Model with a Case Study from Ancient Greek
by: Zafar, Schyan, et al.
Published: (2023)
by: Zafar, Schyan, et al.
Published: (2023)
Opera Graeca Adnotata: Building a 34M+ Token Multilayer Corpus for Ancient Greek
by: Celano, Giuseppe G. A.
Published: (2024)
by: Celano, Giuseppe G. A.
Published: (2024)
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
by: Kou, Qian, et al.
Published: (2026)
by: Kou, Qian, et al.
Published: (2026)
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025)
by: Luo, Jiabin, et al.
Published: (2025)
DOSE: Data Selection for Multi-Modal LLMs via Off-the-Shelf Models
by: Wu, Biao, et al.
Published: (2026)
by: Wu, Biao, et al.
Published: (2026)
PyPotteryLens: An Open-Source Deep Learning Framework for Automated Digitisation of Archaeological Pottery Documentation
by: Cardarelli, Lorenzo
Published: (2024)
by: Cardarelli, Lorenzo
Published: (2024)
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
by: Karamolegkou, Antonia, et al.
Published: (2026)
by: Karamolegkou, Antonia, et al.
Published: (2026)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Instruct-Tuning Pretrained Causal Language Models for Ancient Greek Papyrology and Epigraphy
by: Cullhed, Eric
Published: (2024)
by: Cullhed, Eric
Published: (2024)
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
by: Ma, Dongsheng, et al.
Published: (2026)
by: Ma, Dongsheng, et al.
Published: (2026)
GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek
by: Zhang, Yang, et al.
Published: (2026)
by: Zhang, Yang, et al.
Published: (2026)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
by: Wu, Wei, et al.
Published: (2026)
by: Wu, Wei, et al.
Published: (2026)
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models
by: Ge, Qihang, et al.
Published: (2024)
by: Ge, Qihang, et al.
Published: (2024)
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
Monte Carlo Planning with Large Language Model for Text-Based Game Agents
by: Shi, Zijing, et al.
Published: (2025)
by: Shi, Zijing, et al.
Published: (2025)
Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance
by: Peng, Xueqing, et al.
Published: (2025)
by: Peng, Xueqing, et al.
Published: (2025)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
by: Karim, A H M Rezaul, et al.
Published: (2025)
by: Karim, A H M Rezaul, et al.
Published: (2025)
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Similar Items
-
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
by: Zhang, Nonghai, et al.
Published: (2025) -
Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
by: Yan, Junfeng, et al.
Published: (2025) -
PresentAgent: Multimodal Agent for Presentation Video Generation
by: Shi, Jingwei, et al.
Published: (2025) -
Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
by: Zhang, Zeyu, et al.
Published: (2024) -
Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models
by: Mavromatis, Spyridon, et al.
Published: (2026)