A Framework for Generating Semantically Ambiguous Images to Probe Human and Machine Perception

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yuqi, DuTell, Vasha, Girshick, Ahna R., Corbett, Jennifer E.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915891625590784
author Hu, Yuqi
DuTell, Vasha
Girshick, Ahna R.
Corbett, Jennifer E.
author_facet Hu, Yuqi
DuTell, Vasha
Girshick, Ahna R.
Corbett, Jennifer E.
contents The classic duck-rabbit illusion reveals that when visual evidence is ambiguous, the human brain must decide what it sees. But where exactly do human observers draw the line between ''duck'' and ''rabbit'', and do machine classifiers draw it in the same place? We use semantically ambiguous images as interpretability probes to expose how vision models represent the boundaries between concepts. We present a psychophysically-informed framework that interpolates between concepts in the CLIP embedding space to generate continuous spectra of ambiguous images, allowing us to precisely measure where and how humans and machine classifiers place their semantic boundaries. Using this framework, we show that machine classifiers are more biased towards seeing ''rabbit'', whereas humans are more aligned with the CLIP embedding used for synthesis, and the guidance scale seems to affect human sensitivity more strongly than machine classifiers. Our framework demonstrates how controlled ambiguity can serve as a diagnostic tool to bridge the gap between human psychophysical analysis, image classification, and generative image models, offering insight into human-model alignment, robustness, model interpretability, and image synthesis methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24730
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Framework for Generating Semantically Ambiguous Images to Probe Human and Machine Perception
Hu, Yuqi
DuTell, Vasha
Girshick, Ahna R.
Corbett, Jennifer E.
Computer Vision and Pattern Recognition
The classic duck-rabbit illusion reveals that when visual evidence is ambiguous, the human brain must decide what it sees. But where exactly do human observers draw the line between ''duck'' and ''rabbit'', and do machine classifiers draw it in the same place? We use semantically ambiguous images as interpretability probes to expose how vision models represent the boundaries between concepts. We present a psychophysically-informed framework that interpolates between concepts in the CLIP embedding space to generate continuous spectra of ambiguous images, allowing us to precisely measure where and how humans and machine classifiers place their semantic boundaries. Using this framework, we show that machine classifiers are more biased towards seeing ''rabbit'', whereas humans are more aligned with the CLIP embedding used for synthesis, and the guidance scale seems to affect human sensitivity more strongly than machine classifiers. Our framework demonstrates how controlled ambiguity can serve as a diagnostic tool to bridge the gap between human psychophysical analysis, image classification, and generative image models, offering insight into human-model alignment, robustness, model interpretability, and image synthesis methods.
title A Framework for Generating Semantically Ambiguous Images to Probe Human and Machine Perception
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.24730