Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Jie, Wang, Zhongqi, Lei, Mengqi, Yuan, Zheng, Yan, Bei, Shan, Shiguang, Chen, Xilin
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929728231833600
author Zhang, Jie
Wang, Zhongqi
Lei, Mengqi
Yuan, Zheng
Yan, Bei
Shan, Shiguang
Chen, Xilin
author_facet Zhang, Jie
Wang, Zhongqi
Lei, Mengqi
Yuan, Zheng
Yan, Bei
Shan, Shiguang
Chen, Xilin
contents Currently many benchmarks have been proposed to evaluate the perception ability of the Large Vision-Language Models (LVLMs). However, most benchmarks conduct questions by selecting images from existing datasets, resulting in the potential data leakage. Besides, these benchmarks merely focus on evaluating LVLMs on the realistic style images and clean scenarios, leaving the multi-stylized images and noisy scenarios unexplored. In response to these challenges, we propose a dynamic and scalable benchmark named Dysca for evaluating LVLMs by leveraging synthesis images. Specifically, we leverage Stable Diffusion and design a rule-based method to dynamically generate novel images, questions and the corresponding answers. We consider 51 kinds of image styles and evaluate the perception capability in 20 subtasks. Moreover, we conduct evaluations under 4 scenarios (i.e., Clean, Corruption, Print Attacking and Adversarial Attacking) and 3 question types (i.e., Multi-choices, True-or-false and Free-form). Thanks to the generative paradigm, Dysca serves as a scalable benchmark for easily adding new subtasks and scenarios. A total of 24 advanced open-source LVLMs and 2 close-source LVLMs are evaluated on Dysca, revealing the drawbacks of current LVLMs. The benchmark is released at https://github.com/Robin-WZQ/Dysca.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18849
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
Zhang, Jie
Wang, Zhongqi
Lei, Mengqi
Yuan, Zheng
Yan, Bei
Shan, Shiguang
Chen, Xilin
Computer Vision and Pattern Recognition
Currently many benchmarks have been proposed to evaluate the perception ability of the Large Vision-Language Models (LVLMs). However, most benchmarks conduct questions by selecting images from existing datasets, resulting in the potential data leakage. Besides, these benchmarks merely focus on evaluating LVLMs on the realistic style images and clean scenarios, leaving the multi-stylized images and noisy scenarios unexplored. In response to these challenges, we propose a dynamic and scalable benchmark named Dysca for evaluating LVLMs by leveraging synthesis images. Specifically, we leverage Stable Diffusion and design a rule-based method to dynamically generate novel images, questions and the corresponding answers. We consider 51 kinds of image styles and evaluate the perception capability in 20 subtasks. Moreover, we conduct evaluations under 4 scenarios (i.e., Clean, Corruption, Print Attacking and Adversarial Attacking) and 3 question types (i.e., Multi-choices, True-or-false and Free-form). Thanks to the generative paradigm, Dysca serves as a scalable benchmark for easily adding new subtasks and scenarios. A total of 24 advanced open-source LVLMs and 2 close-source LVLMs are evaluated on Dysca, revealing the drawbacks of current LVLMs. The benchmark is released at https://github.com/Robin-WZQ/Dysca.
title Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.18849