Saved in:
Bibliographic Details
Main Authors: Zhang, Qing, Li, Xuesong, Zhang, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.20501
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915863780655104
author Zhang, Qing
Li, Xuesong
Zhang, Jing
author_facet Zhang, Qing
Li, Xuesong
Zhang, Jing
contents What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction, and interaction perception, which models how an agent's actions engage with those parts. To test this hypothesis, we conduct a systematic probing of Visual Foundation Models (VFMs). We find that models like DINO inherently encode part-level geometric structures, while generative models like Flux contain rich, verb-conditioned spatial attention maps that serve as implicit interaction priors. Crucially, we demonstrate that these two dimensions are not merely correlated but are composable elements of affordance. By simply fusing DINO's geometric prototypes with Flux's interaction maps in a training-free and zero-shot manner, we achieve affordance estimation competitive with weakly-supervised methods. This final fusion experiment confirms that geometric and interaction perception are the fundamental building blocks of affordance understanding in VFMs, providing a mechanistic account of how perception grounds action.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20501
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models
Zhang, Qing
Li, Xuesong
Zhang, Jing
Computer Vision and Pattern Recognition
What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction, and interaction perception, which models how an agent's actions engage with those parts. To test this hypothesis, we conduct a systematic probing of Visual Foundation Models (VFMs). We find that models like DINO inherently encode part-level geometric structures, while generative models like Flux contain rich, verb-conditioned spatial attention maps that serve as implicit interaction priors. Crucially, we demonstrate that these two dimensions are not merely correlated but are composable elements of affordance. By simply fusing DINO's geometric prototypes with Flux's interaction maps in a training-free and zero-shot manner, we achieve affordance estimation competitive with weakly-supervised methods. This final fusion experiment confirms that geometric and interaction perception are the fundamental building blocks of affordance understanding in VFMs, providing a mechanistic account of how perception grounds action.
title Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.20501