Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuhui, Su, Yuchang, Liu, Yiming, Wang, Xiaohan, Burgess, James, Sui, Elaine, Wang, Chenyu, Aklilu, Josiah, Lozano, Alejandro, Wei, Anjiang, Schmidt, Ludwig, Yeung-Levy, Serena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913785534480384
author Zhang, Yuhui
Su, Yuchang
Liu, Yiming
Wang, Xiaohan
Burgess, James
Sui, Elaine
Wang, Chenyu
Aklilu, Josiah
Lozano, Alejandro
Wei, Anjiang
Schmidt, Ludwig
Yeung-Levy, Serena
author_facet Zhang, Yuhui
Su, Yuchang
Liu, Yiming
Wang, Xiaohan
Burgess, James
Sui, Elaine
Wang, Chenyu
Aklilu, Josiah
Lozano, Alejandro
Wei, Anjiang
Schmidt, Ludwig
Yeung-Levy, Serena
contents The rapid development of vision language models (VLMs) demands rigorous and reliable evaluation. However, current visual question answering (VQA) benchmarks often depend on open-ended questions, making accurate evaluation difficult due to the variability in natural language responses. To address this, we introduce AutoConverter, an agentic framework that automatically converts these open-ended questions into multiple-choice format, enabling objective evaluation while reducing the costly multiple-choice question creation process. Our experiments demonstrate that AutoConverter can generate correct and challenging multiple-choice questions, with VLMs demonstrating consistently similar or lower accuracy on these questions compared to human-created ones. Using AutoConverter, we construct VMCBench, a benchmark created by transforming 20 existing VQA datasets into a unified multiple-choice format, totaling 9,018 questions. We comprehensively evaluate 33 state-of-the-art VLMs on VMCBench, setting a new standard for scalable, consistent, and reproducible VLM evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03225
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
Zhang, Yuhui
Su, Yuchang
Liu, Yiming
Wang, Xiaohan
Burgess, James
Sui, Elaine
Wang, Chenyu
Aklilu, Josiah
Lozano, Alejandro
Wei, Anjiang
Schmidt, Ludwig
Yeung-Levy, Serena
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
The rapid development of vision language models (VLMs) demands rigorous and reliable evaluation. However, current visual question answering (VQA) benchmarks often depend on open-ended questions, making accurate evaluation difficult due to the variability in natural language responses. To address this, we introduce AutoConverter, an agentic framework that automatically converts these open-ended questions into multiple-choice format, enabling objective evaluation while reducing the costly multiple-choice question creation process. Our experiments demonstrate that AutoConverter can generate correct and challenging multiple-choice questions, with VLMs demonstrating consistently similar or lower accuracy on these questions compared to human-created ones. Using AutoConverter, we construct VMCBench, a benchmark created by transforming 20 existing VQA datasets into a unified multiple-choice format, totaling 9,018 questions. We comprehensively evaluate 33 state-of-the-art VLMs on VMCBench, setting a new standard for scalable, consistent, and reproducible VLM evaluation.
title Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2501.03225