Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rostamkhani, Mohammadmostafa, Ansari, Baktash, Sabzevari, Hoorieh, Rahmani, Farzan, Eetemadi, Sauleh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912152124653568
author Rostamkhani, Mohammadmostafa
Ansari, Baktash
Sabzevari, Hoorieh
Rahmani, Farzan
Eetemadi, Sauleh
author_facet Rostamkhani, Mohammadmostafa
Ansari, Baktash
Sabzevari, Hoorieh
Rahmani, Farzan
Eetemadi, Sauleh
contents In recent years, Visual Question Answering (VQA) has made significant strides, particularly with the advent of multimodal models that integrate vision and language understanding. However, existing VQA datasets often overlook the complexities introduced by image illusions, which pose unique challenges for both human perception and model interpretation. In this study, we introduce a novel task called Illusory VQA, along with four specialized datasets: IllusionMNIST, IllusionFashionMNIST, IllusionAnimals, and IllusionChar. These datasets are designed to evaluate the performance of state-of-the-art multimodal models in recognizing and interpreting visual illusions. We assess the zero-shot performance of various models, fine-tune selected models on our datasets, and propose a simple yet effective solution for illusion detection using Gaussian and blur low-pass filters. We show that this method increases the performance of models significantly and in the case of BLIP-2 on IllusionAnimals without any fine-tuning, it outperforms humans. Our findings highlight the disparity between human and model perception of illusions and demonstrate that fine-tuning and specific preprocessing techniques can significantly enhance model robustness. This work contributes to the development of more human-like visual understanding in multimodal models and suggests future directions for adapting filters using learnable parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08169
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
Rostamkhani, Mohammadmostafa
Ansari, Baktash
Sabzevari, Hoorieh
Rahmani, Farzan
Eetemadi, Sauleh
Computer Vision and Pattern Recognition
Computation and Language
In recent years, Visual Question Answering (VQA) has made significant strides, particularly with the advent of multimodal models that integrate vision and language understanding. However, existing VQA datasets often overlook the complexities introduced by image illusions, which pose unique challenges for both human perception and model interpretation. In this study, we introduce a novel task called Illusory VQA, along with four specialized datasets: IllusionMNIST, IllusionFashionMNIST, IllusionAnimals, and IllusionChar. These datasets are designed to evaluate the performance of state-of-the-art multimodal models in recognizing and interpreting visual illusions. We assess the zero-shot performance of various models, fine-tune selected models on our datasets, and propose a simple yet effective solution for illusion detection using Gaussian and blur low-pass filters. We show that this method increases the performance of models significantly and in the case of BLIP-2 on IllusionAnimals without any fine-tuning, it outperforms humans. Our findings highlight the disparity between human and model perception of illusions and demonstrate that fine-tuning and specific preprocessing techniques can significantly enhance model robustness. This work contributes to the development of more human-like visual understanding in multimodal models and suggests future directions for adapting filters using learnable parameters.
title Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2412.08169