VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916038124240896 |
|---|---|
| author | Li, Zhaonan Chickering, Kyle R. Li, Bangzheng Dineen, Jacob Ye, Xiao Xu, Zhikun Lu, Shijie Huang, Yuxi Shen, Ming Nguyen, Bach Pavuluri, Jaya Adithya Nguyen, Mau Son Chavan, Sanika Le, Ngoc Minh Thu Chen, Muhao Zhou, Ben |
| author_facet | Li, Zhaonan Chickering, Kyle R. Li, Bangzheng Dineen, Jacob Ye, Xiao Xu, Zhikun Lu, Shijie Huang, Yuxi Shen, Ming Nguyen, Bach Pavuluri, Jaya Adithya Nguyen, Mau Son Chavan, Sanika Le, Ngoc Minh Thu Chen, Muhao Zhou, Ben |
| contents | A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipulate concept-level properties under transformation and transfer them to new scenes. We introduce VisAnalog, a controlled suite for this setting on natural images. Each example instantiates $A\!:\!B::C\!:\,?$: images $B$ and a hidden target image $D$ are produced by applying the same deterministic transformation sequence to source images $A$ and $C$. Given $A$, $B$, and $C$, a model must answer a multiple-choice question about $D$. The benchmark contains 617 human-validated questions spanning one- to four-step transformations such as zoom, quadrant swap, rotation, flip, and hue rotation. Across strong proprietary and open-source VLMs, end-to-end accuracy is substantially lower than oracle accuracy when $D$ is directly shown, and degrades sharply as transformation depth increases, while human performance remains near the ceiling. A program-conditioned evaluation further separates failures of relation inference from failures of transformation application, showing that inferring the visual relation from $A \rightarrow B$ is the dominant bottleneck, with additional application errors emerging on harder multi-step cases. The dataset is publicly available at https://huggingface.co/datasets/zli99/VisAnalog. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_23141 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images Li, Zhaonan Chickering, Kyle R. Li, Bangzheng Dineen, Jacob Ye, Xiao Xu, Zhikun Lu, Shijie Huang, Yuxi Shen, Ming Nguyen, Bach Pavuluri, Jaya Adithya Nguyen, Mau Son Chavan, Sanika Le, Ngoc Minh Thu Chen, Muhao Zhou, Ben Computer Vision and Pattern Recognition A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipulate concept-level properties under transformation and transfer them to new scenes. We introduce VisAnalog, a controlled suite for this setting on natural images. Each example instantiates $A\!:\!B::C\!:\,?$: images $B$ and a hidden target image $D$ are produced by applying the same deterministic transformation sequence to source images $A$ and $C$. Given $A$, $B$, and $C$, a model must answer a multiple-choice question about $D$. The benchmark contains 617 human-validated questions spanning one- to four-step transformations such as zoom, quadrant swap, rotation, flip, and hue rotation. Across strong proprietary and open-source VLMs, end-to-end accuracy is substantially lower than oracle accuracy when $D$ is directly shown, and degrades sharply as transformation depth increases, while human performance remains near the ceiling. A program-conditioned evaluation further separates failures of relation inference from failures of transformation application, showing that inferring the visual relation from $A \rightarrow B$ is the dominant bottleneck, with additional application errors emerging on harder multi-step cases. The dataset is publicly available at https://huggingface.co/datasets/zli99/VisAnalog. |
| title | VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2605.23141 |