Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gaur, Manu, S, Darshan Singh, Tapaswi, Makarand
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914955257708544
author Gaur, Manu
S, Darshan Singh
Tapaswi, Makarand
author_facet Gaur, Manu
S, Darshan Singh
Tapaswi, Makarand
contents Visual Question Answering (VQA) with multiple choice questions enables a vision-centric evaluation of Multimodal Large Language Models (MLLMs). Although it reliably checks the existence of specific visual abilities, it is easier for the model to select an answer from multiple choices (VQA evaluation) than to generate the answer itself. In this work, we offer a novel perspective: we evaluate how well an MLLM understands a specific visual concept by its ability to uniquely describe two extremely similar images that differ only in the targeted visual concept. Specifically, we assess the ability of MLLMs to capture specific points of visual differences using self-retrieval, i.e., by retrieving the target image using its generated caption against the other image in the pair serving as the distractor. We curate 247 highly similar image pairs as part of the D3 benchmark. For each image pair, the model is prompted to: (1) Detect a specific visual difference, and (2) Describe the target image uniquely such that it (3) Discriminates the target image from the distractor. Self-retrieval within D3 enables whitebox evaluation across six different visual patterns, revealing that current models struggle to independently discern fine-grained visual differences, with open-source models failing to outperform random guess.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15125
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
Gaur, Manu
S, Darshan Singh
Tapaswi, Makarand
Computer Vision and Pattern Recognition
Visual Question Answering (VQA) with multiple choice questions enables a vision-centric evaluation of Multimodal Large Language Models (MLLMs). Although it reliably checks the existence of specific visual abilities, it is easier for the model to select an answer from multiple choices (VQA evaluation) than to generate the answer itself. In this work, we offer a novel perspective: we evaluate how well an MLLM understands a specific visual concept by its ability to uniquely describe two extremely similar images that differ only in the targeted visual concept. Specifically, we assess the ability of MLLMs to capture specific points of visual differences using self-retrieval, i.e., by retrieving the target image using its generated caption against the other image in the pair serving as the distractor. We curate 247 highly similar image pairs as part of the D3 benchmark. For each image pair, the model is prompted to: (1) Detect a specific visual difference, and (2) Describe the target image uniquely such that it (3) Discriminates the target image from the distractor. Self-retrieval within D3 enables whitebox evaluation across six different visual patterns, revealing that current models struggle to independently discern fine-grained visual differences, with open-source models failing to outperform random guess.
title Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.15125