VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Yanling, Zhao, Yihan, Chen, Xiaodong, Guo, Shasha, Liu, Lixin, Li, Haoyang, Xiao, Yong, Zhang, Jing, Li, Qi, Xu, Ke
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913725743628288
author Wang, Yanling
Zhao, Yihan
Chen, Xiaodong
Guo, Shasha
Liu, Lixin
Li, Haoyang
Xiao, Yong
Zhang, Jing
Li, Qi
Xu, Ke
author_facet Wang, Yanling
Zhao, Yihan
Chen, Xiaodong
Guo, Shasha
Liu, Lixin
Li, Haoyang
Xiao, Yong
Zhang, Jing
Li, Qi
Xu, Ke
contents Large vision-language models (LVLMs) have demonstrated remarkable achievements, yet the generation of non-factual responses remains prevalent in fact-seeking question answering (QA). Current multimodal fact-seeking benchmarks primarily focus on comparing model outputs to ground truth answers, providing limited insights into the performance of modality-specific modules. To bridge this gap, we introduce VisualSimpleQA, a multimodal fact-seeking benchmark with two key features. First, it enables streamlined and decoupled evaluation of LVLMs in visual and linguistic modalities. Second, it incorporates well-defined difficulty criteria to guide human annotation and facilitates the extraction of a challenging subset, VisualSimpleQA-hard. Experiments on 15 LVLMs show that even state-of-the-art models such as GPT-4o achieve merely 60%+ correctness in multimodal fact-seeking QA on VisualSimpleQA and 30%+ on VisualSimpleQA-hard. Furthermore, the decoupled evaluation across these models highlights substantial opportunities for improvement in both visual and linguistic modules. The dataset is available at https://huggingface.co/datasets/WYLing/VisualSimpleQA.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
Wang, Yanling
Zhao, Yihan
Chen, Xiaodong
Guo, Shasha
Liu, Lixin
Li, Haoyang
Xiao, Yong
Zhang, Jing
Li, Qi
Xu, Ke
Computation and Language
Computer Vision and Pattern Recognition
Large vision-language models (LVLMs) have demonstrated remarkable achievements, yet the generation of non-factual responses remains prevalent in fact-seeking question answering (QA). Current multimodal fact-seeking benchmarks primarily focus on comparing model outputs to ground truth answers, providing limited insights into the performance of modality-specific modules. To bridge this gap, we introduce VisualSimpleQA, a multimodal fact-seeking benchmark with two key features. First, it enables streamlined and decoupled evaluation of LVLMs in visual and linguistic modalities. Second, it incorporates well-defined difficulty criteria to guide human annotation and facilitates the extraction of a challenging subset, VisualSimpleQA-hard. Experiments on 15 LVLMs show that even state-of-the-art models such as GPT-4o achieve merely 60%+ correctness in multimodal fact-seeking QA on VisualSimpleQA and 30%+ on VisualSimpleQA-hard. Furthermore, the decoupled evaluation across these models highlights substantial opportunities for improvement in both visual and linguistic modules. The dataset is available at https://huggingface.co/datasets/WYLing/VisualSimpleQA.
title VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06492