Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913785534480384 |
|---|---|
| author | Zhang, Yuhui Su, Yuchang Liu, Yiming Wang, Xiaohan Burgess, James Sui, Elaine Wang, Chenyu Aklilu, Josiah Lozano, Alejandro Wei, Anjiang Schmidt, Ludwig Yeung-Levy, Serena |
| author_facet | Zhang, Yuhui Su, Yuchang Liu, Yiming Wang, Xiaohan Burgess, James Sui, Elaine Wang, Chenyu Aklilu, Josiah Lozano, Alejandro Wei, Anjiang Schmidt, Ludwig Yeung-Levy, Serena |
| contents | The rapid development of vision language models (VLMs) demands rigorous and reliable evaluation. However, current visual question answering (VQA) benchmarks often depend on open-ended questions, making accurate evaluation difficult due to the variability in natural language responses. To address this, we introduce AutoConverter, an agentic framework that automatically converts these open-ended questions into multiple-choice format, enabling objective evaluation while reducing the costly multiple-choice question creation process. Our experiments demonstrate that AutoConverter can generate correct and challenging multiple-choice questions, with VLMs demonstrating consistently similar or lower accuracy on these questions compared to human-created ones. Using AutoConverter, we construct VMCBench, a benchmark created by transforming 20 existing VQA datasets into a unified multiple-choice format, totaling 9,018 questions. We comprehensively evaluate 33 state-of-the-art VLMs on VMCBench, setting a new standard for scalable, consistent, and reproducible VLM evaluation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_03225 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation Zhang, Yuhui Su, Yuchang Liu, Yiming Wang, Xiaohan Burgess, James Sui, Elaine Wang, Chenyu Aklilu, Josiah Lozano, Alejandro Wei, Anjiang Schmidt, Ludwig Yeung-Levy, Serena Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Computers and Society Machine Learning The rapid development of vision language models (VLMs) demands rigorous and reliable evaluation. However, current visual question answering (VQA) benchmarks often depend on open-ended questions, making accurate evaluation difficult due to the variability in natural language responses. To address this, we introduce AutoConverter, an agentic framework that automatically converts these open-ended questions into multiple-choice format, enabling objective evaluation while reducing the costly multiple-choice question creation process. Our experiments demonstrate that AutoConverter can generate correct and challenging multiple-choice questions, with VLMs demonstrating consistently similar or lower accuracy on these questions compared to human-created ones. Using AutoConverter, we construct VMCBench, a benchmark created by transforming 20 existing VQA datasets into a unified multiple-choice format, totaling 9,018 questions. We comprehensively evaluate 33 state-of-the-art VLMs on VMCBench, setting a new standard for scalable, consistent, and reproducible VLM evaluation. |
| title | Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Computers and Society Machine Learning |
| url | https://arxiv.org/abs/2501.03225 |