AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bao, Han, Huang, Yue, Wang, Yanbo, Ye, Jiayi, Wang, Xiangqi, Chen, Xiuying, Zhao, Yue, Zhou, Tianyi, Elhoseiny, Mohamed, Zhang, Xiangliang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929743850373120
author Bao, Han
Huang, Yue
Wang, Yanbo
Ye, Jiayi
Wang, Xiangqi
Chen, Xiuying
Zhao, Yue
Zhou, Tianyi
Elhoseiny, Mohamed
Zhang, Xiangliang
author_facet Bao, Han
Huang, Yue
Wang, Yanbo
Ye, Jiayi
Wang, Xiangqi
Chen, Xiuying
Zhao, Yue
Zhou, Tianyi
Elhoseiny, Mohamed
Zhang, Xiangliang
contents Large Vision-Language Models (LVLMs) have become essential for advancing the integration of visual and linguistic information. However, the evaluation of LVLMs presents significant challenges as the evaluation benchmark always demands lots of human cost for its construction, and remains static, lacking flexibility once constructed. Even though automatic evaluation has been explored in textual modality, the visual modality remains under-explored. As a result, in this work, we address a question: "Can LVLMs themselves be used to benchmark each other in the visual automatically domain?". We introduce AutoBench-V, an automated framework for serving evaluation on demand, i.e., benchmarking LVLMs based on specific aspects of model capability. AutoBench-V leverages text-to-image models to generate relevant image samples and then utilizes LVLMs to orchestrate visual question-answering (VQA) tasks, completing the evaluation process efficiently and flexibly. Through an extensive evaluation of nine popular LVLMs across five demanded user inputs (i.e., evaluation capabilities), the framework shows effectiveness and reliability.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21259
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
Bao, Han
Huang, Yue
Wang, Yanbo
Ye, Jiayi
Wang, Xiangqi
Chen, Xiuying
Zhao, Yue
Zhou, Tianyi
Elhoseiny, Mohamed
Zhang, Xiangliang
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Vision-Language Models (LVLMs) have become essential for advancing the integration of visual and linguistic information. However, the evaluation of LVLMs presents significant challenges as the evaluation benchmark always demands lots of human cost for its construction, and remains static, lacking flexibility once constructed. Even though automatic evaluation has been explored in textual modality, the visual modality remains under-explored. As a result, in this work, we address a question: "Can LVLMs themselves be used to benchmark each other in the visual automatically domain?". We introduce AutoBench-V, an automated framework for serving evaluation on demand, i.e., benchmarking LVLMs based on specific aspects of model capability. AutoBench-V leverages text-to-image models to generate relevant image samples and then utilizes LVLMs to orchestrate visual question-answering (VQA) tasks, completing the evaluation process efficiently and flexibly. Through an extensive evaluation of nine popular LVLMs across five demanded user inputs (i.e., evaluation capabilities), the framework shows effectiveness and reliability.
title AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2410.21259