QuizRank: Picking Images by Quizzing VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Tenghao, Adar, Eytan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911165008838656
author Ji, Tenghao
Adar, Eytan
author_facet Ji, Tenghao
Adar, Eytan
contents Images play a vital role in improving the readability and comprehension of Wikipedia articles by serving as `illustrative aids.' However, not all images are equally effective and not all Wikipedia editors are trained in their selection. We propose QuizRank, a novel method of image selection that leverages large language models (LLMs) and vision language models (VLMs) to rank images as learning interventions. Our approach transforms textual descriptions of the article's subject into multiple-choice questions about important visual characteristics of the concept. We utilize these questions to quiz the VLM: the better an image can help answer questions, the higher it is ranked. To further improve discrimination between visually similar items, we introduce a Contrastive QuizRank that leverages differences in the features of target (e.g., a Western Bluebird) and distractor concepts (e.g., Mountain Bluebird) to generate questions. We demonstrate the potential of VLMs as effective visual evaluators by showing a high congruence with human quiz-takers and an effective discriminative ranking of images.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15059
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle QuizRank: Picking Images by Quizzing VLMs
Ji, Tenghao
Adar, Eytan
Human-Computer Interaction
Computer Vision and Pattern Recognition
Images play a vital role in improving the readability and comprehension of Wikipedia articles by serving as `illustrative aids.' However, not all images are equally effective and not all Wikipedia editors are trained in their selection. We propose QuizRank, a novel method of image selection that leverages large language models (LLMs) and vision language models (VLMs) to rank images as learning interventions. Our approach transforms textual descriptions of the article's subject into multiple-choice questions about important visual characteristics of the concept. We utilize these questions to quiz the VLM: the better an image can help answer questions, the higher it is ranked. To further improve discrimination between visually similar items, we introduce a Contrastive QuizRank that leverages differences in the features of target (e.g., a Western Bluebird) and distractor concepts (e.g., Mountain Bluebird) to generate questions. We demonstrate the potential of VLMs as effective visual evaluators by showing a high congruence with human quiz-takers and an effective discriminative ranking of images.
title QuizRank: Picking Images by Quizzing VLMs
topic Human-Computer Interaction
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.15059