AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vu, Nghia, Do, Tuong, Nguyen, Khang, Huang, Baoru, Le, Nhat, Nguyen, Binh Xuan, Tjiputra, Erman, Tran, Quang D., Prakash, Ravi, Chiu, Te-Chuan, Nguyen, Anh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914431456247808
author Vu, Nghia
Do, Tuong
Nguyen, Khang
Huang, Baoru
Le, Nhat
Nguyen, Binh Xuan
Tjiputra, Erman
Tran, Quang D.
Prakash, Ravi
Chiu, Te-Chuan
Nguyen, Anh
author_facet Vu, Nghia
Do, Tuong
Nguyen, Khang
Huang, Baoru
Le, Nhat
Nguyen, Binh Xuan
Tjiputra, Erman
Tran, Quang D.
Prakash, Ravi
Chiu, Te-Chuan
Nguyen, Anh
contents Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of objects to determine interactable regions. However, extending this learning capability to a scene is significantly more complicated, as incorporating object- and scene-level semantics is not straightforward. In this work, we introduce AffordBridge, a large-scale dataset with 291,637 functional interaction annotations across 685 high-resolution indoor scenes in the form of point clouds. Our affordance annotations are complemented by RGB images that are linked to the same instances within the scenes. Building upon our dataset, we propose AffordMatcher, an affordance learning method that establishes coherent semantic correspondences between image-based and point cloud-based instances for keypoint matching, enabling a more precise identification of affordance regions based on cues, so-called visual signifiers. Experimental results on our dataset demonstrate the effectiveness of our approach compared to other methods.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27970
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
Vu, Nghia
Do, Tuong
Nguyen, Khang
Huang, Baoru
Le, Nhat
Nguyen, Binh Xuan
Tjiputra, Erman
Tran, Quang D.
Prakash, Ravi
Chiu, Te-Chuan
Nguyen, Anh
Computer Vision and Pattern Recognition
Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of objects to determine interactable regions. However, extending this learning capability to a scene is significantly more complicated, as incorporating object- and scene-level semantics is not straightforward. In this work, we introduce AffordBridge, a large-scale dataset with 291,637 functional interaction annotations across 685 high-resolution indoor scenes in the form of point clouds. Our affordance annotations are complemented by RGB images that are linked to the same instances within the scenes. Building upon our dataset, we propose AffordMatcher, an affordance learning method that establishes coherent semantic correspondences between image-based and point cloud-based instances for keypoint matching, enabling a more precise identification of affordance regions based on cues, so-called visual signifiers. Experimental results on our dataset demonstrate the effectiveness of our approach compared to other methods.
title AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.27970