The Illusion-Illusion: Vision Language Models See Illusions Where There are None

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Ullman, Tomer
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929646665203712
author Ullman, Tomer
author_facet Ullman, Tomer
contents Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something "really is" and how something "appears to be", and this gap helps us understand the mental processing that lead to how something appears to be. Illusions are also useful for investigating artificial systems, and much research has examined whether computational models of perceptions fall prey to the same illusions as people. Here, I invert the standard use of perceptual illusions to examine basic processing errors in current vision language models. I present these models with illusory-illusions, neighbors of common illusions that should not elicit processing errors. These include such things as perfectly reasonable ducks, crooked lines that truly are crooked, circles that seem to have different sizes because they are, in fact, of different sizes, and so on. I show that many current vision language systems mistakenly see these illusion-illusions as illusions. I suggest that such failures are part of broader failures already discussed in the literature.
format Preprint
id arxiv_https___arxiv_org_abs_2412_18613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Illusion-Illusion: Vision Language Models See Illusions Where There are None
Ullman, Tomer
Neurons and Cognition
Computation and Language
Computer Vision and Pattern Recognition
Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something "really is" and how something "appears to be", and this gap helps us understand the mental processing that lead to how something appears to be. Illusions are also useful for investigating artificial systems, and much research has examined whether computational models of perceptions fall prey to the same illusions as people. Here, I invert the standard use of perceptual illusions to examine basic processing errors in current vision language models. I present these models with illusory-illusions, neighbors of common illusions that should not elicit processing errors. These include such things as perfectly reasonable ducks, crooked lines that truly are crooked, circles that seem to have different sizes because they are, in fact, of different sizes, and so on. I show that many current vision language systems mistakenly see these illusion-illusions as illusions. I suggest that such failures are part of broader failures already discussed in the literature.
title The Illusion-Illusion: Vision Language Models See Illusions Where There are None
topic Neurons and Cognition
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.18613