VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Zhaonan, Chickering, Kyle R., Li, Bangzheng, Dineen, Jacob, Ye, Xiao, Xu, Zhikun, Lu, Shijie, Huang, Yuxi, Shen, Ming, Nguyen, Bach, Pavuluri, Jaya Adithya, Nguyen, Mau Son, Chavan, Sanika, Le, Ngoc Minh Thu, Chen, Muhao, Zhou, Ben
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916038124240896
author Li, Zhaonan
Chickering, Kyle R.
Li, Bangzheng
Dineen, Jacob
Ye, Xiao
Xu, Zhikun
Lu, Shijie
Huang, Yuxi
Shen, Ming
Nguyen, Bach
Pavuluri, Jaya Adithya
Nguyen, Mau Son
Chavan, Sanika
Le, Ngoc Minh Thu
Chen, Muhao
Zhou, Ben
author_facet Li, Zhaonan
Chickering, Kyle R.
Li, Bangzheng
Dineen, Jacob
Ye, Xiao
Xu, Zhikun
Lu, Shijie
Huang, Yuxi
Shen, Ming
Nguyen, Bach
Pavuluri, Jaya Adithya
Nguyen, Mau Son
Chavan, Sanika
Le, Ngoc Minh Thu
Chen, Muhao
Zhou, Ben
contents A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipulate concept-level properties under transformation and transfer them to new scenes. We introduce VisAnalog, a controlled suite for this setting on natural images. Each example instantiates $A\!:\!B::C\!:\,?$: images $B$ and a hidden target image $D$ are produced by applying the same deterministic transformation sequence to source images $A$ and $C$. Given $A$, $B$, and $C$, a model must answer a multiple-choice question about $D$. The benchmark contains 617 human-validated questions spanning one- to four-step transformations such as zoom, quadrant swap, rotation, flip, and hue rotation. Across strong proprietary and open-source VLMs, end-to-end accuracy is substantially lower than oracle accuracy when $D$ is directly shown, and degrades sharply as transformation depth increases, while human performance remains near the ceiling. A program-conditioned evaluation further separates failures of relation inference from failures of transformation application, showing that inferring the visual relation from $A \rightarrow B$ is the dominant bottleneck, with additional application errors emerging on harder multi-step cases. The dataset is publicly available at https://huggingface.co/datasets/zli99/VisAnalog.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23141
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
Li, Zhaonan
Chickering, Kyle R.
Li, Bangzheng
Dineen, Jacob
Ye, Xiao
Xu, Zhikun
Lu, Shijie
Huang, Yuxi
Shen, Ming
Nguyen, Bach
Pavuluri, Jaya Adithya
Nguyen, Mau Son
Chavan, Sanika
Le, Ngoc Minh Thu
Chen, Muhao
Zhou, Ben
Computer Vision and Pattern Recognition
A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipulate concept-level properties under transformation and transfer them to new scenes. We introduce VisAnalog, a controlled suite for this setting on natural images. Each example instantiates $A\!:\!B::C\!:\,?$: images $B$ and a hidden target image $D$ are produced by applying the same deterministic transformation sequence to source images $A$ and $C$. Given $A$, $B$, and $C$, a model must answer a multiple-choice question about $D$. The benchmark contains 617 human-validated questions spanning one- to four-step transformations such as zoom, quadrant swap, rotation, flip, and hue rotation. Across strong proprietary and open-source VLMs, end-to-end accuracy is substantially lower than oracle accuracy when $D$ is directly shown, and degrades sharply as transformation depth increases, while human performance remains near the ceiling. A program-conditioned evaluation further separates failures of relation inference from failures of transformation application, showing that inferring the visual relation from $A \rightarrow B$ is the dominant bottleneck, with additional application errors emerging on harder multi-step cases. The dataset is publicly available at https://huggingface.co/datasets/zli99/VisAnalog.
title VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.23141