UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Qiaojun, Huang, Siyuan, Yuan, Xibin, Jiang, Zhengkai, Hao, Ce, Li, Xin, Chang, Haonan, Wang, Junbo, Liu, Liu, Li, Hongsheng, Gao, Peng, Lu, Cewu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912222940233728
author Yu, Qiaojun
Huang, Siyuan
Yuan, Xibin
Jiang, Zhengkai
Hao, Ce
Li, Xin
Chang, Haonan
Wang, Junbo
Liu, Liu
Li, Hongsheng
Gao, Peng
Lu, Cewu
author_facet Yu, Qiaojun
Huang, Siyuan
Yuan, Xibin
Jiang, Zhengkai
Hao, Ce
Li, Xin
Chang, Haonan
Wang, Junbo
Liu, Liu
Li, Hongsheng
Gao, Peng
Lu, Cewu
contents Previous studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified formulation. Specifically, we constructed a dataset labeled with manipulation-related key attributes, comprising 900 articulated objects from 19 categories and 600 tools from 12 categories. Furthermore, we leverage MLLMs to infer object-centric representations for manipulation tasks, including affordance recognition and reasoning about 3D motion constraints. Comprehensive experiments in both simulation and real-world settings indicate that UniAff significantly improves the generalization of robotic manipulation for tools and articulated objects. We hope that UniAff will serve as a general baseline for unified robotic manipulation tasks in the future. Images, videos, dataset, and code are published on the project website at:https://sites.google.com/view/uni-aff/home
format Preprint
id arxiv_https___arxiv_org_abs_2409_20551
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
Yu, Qiaojun
Huang, Siyuan
Yuan, Xibin
Jiang, Zhengkai
Hao, Ce
Li, Xin
Chang, Haonan
Wang, Junbo
Liu, Liu
Li, Hongsheng
Gao, Peng
Lu, Cewu
Robotics
Previous studies on robotic manipulation are based on a limited understanding of the underlying 3D motion constraints and affordances. To address these challenges, we propose a comprehensive paradigm, termed UniAff, that integrates 3D object-centric manipulation and task understanding in a unified formulation. Specifically, we constructed a dataset labeled with manipulation-related key attributes, comprising 900 articulated objects from 19 categories and 600 tools from 12 categories. Furthermore, we leverage MLLMs to infer object-centric representations for manipulation tasks, including affordance recognition and reasoning about 3D motion constraints. Comprehensive experiments in both simulation and real-world settings indicate that UniAff significantly improves the generalization of robotic manipulation for tools and articulated objects. We hope that UniAff will serve as a general baseline for unified robotic manipulation tasks in the future. Images, videos, dataset, and code are published on the project website at:https://sites.google.com/view/uni-aff/home
title UniAff: A Unified Representation of Affordances for Tool Usage and Articulation with Vision-Language Models
topic Robotics
url https://arxiv.org/abs/2409.20551