IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Can, Lee, Gim Hee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909572771348480
author Zhang, Can
Lee, Gim Hee
author_facet Zhang, Can
Lee, Gim Hee
contents This work presents IAAO, a novel framework that builds an explicit 3D model for intelligent agents to gain understanding of articulated objects in their environment through interaction. Unlike prior methods that rely on task-specific networks and assumptions about movable parts, our IAAO leverages large foundation models to estimate interactive affordances and part articulations in three stages. We first build hierarchical features and label fields for each object state using 3D Gaussian Splatting (3DGS) by distilling mask features and view-consistent labels from multi-view images. We then perform object- and part-level queries on the 3D Gaussian primitives to identify static and articulated elements, estimating global transformations and local articulation parameters along with affordances. Finally, scenes from different states are merged and refined based on the estimated transformations, enabling robust affordance-based interaction and manipulation of objects. Experimental results demonstrate the effectiveness of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06827
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments
Zhang, Can
Lee, Gim Hee
Computer Vision and Pattern Recognition
This work presents IAAO, a novel framework that builds an explicit 3D model for intelligent agents to gain understanding of articulated objects in their environment through interaction. Unlike prior methods that rely on task-specific networks and assumptions about movable parts, our IAAO leverages large foundation models to estimate interactive affordances and part articulations in three stages. We first build hierarchical features and label fields for each object state using 3D Gaussian Splatting (3DGS) by distilling mask features and view-consistent labels from multi-view images. We then perform object- and part-level queries on the 3D Gaussian primitives to identify static and articulated elements, estimating global transformations and local articulation parameters along with affordances. Finally, scenes from different states are merged and refined based on the estimated transformations, enabling robust affordance-based interaction and manipulation of objects. Experimental results demonstrate the effectiveness of our method.
title IAAO: Interactive Affordance Learning for Articulated Objects in 3D Environments
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.06827