3D Foundation Models Enable Simultaneous Geometry and Pose Estimation of Grasped Objects

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhi, Weiming, Tang, Haozhan, Zhang, Tianyi, Johnson-Roberson, Matthew
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910527413813248
author Zhi, Weiming
Tang, Haozhan
Zhang, Tianyi
Johnson-Roberson, Matthew
author_facet Zhi, Weiming
Tang, Haozhan
Zhang, Tianyi
Johnson-Roberson, Matthew
contents Humans have the remarkable ability to use held objects as tools to interact with their environment. For this to occur, humans internally estimate how hand movements affect the object's movement. We wish to endow robots with this capability. We contribute methodology to jointly estimate the geometry and pose of objects grasped by a robot, from RGB images captured by an external camera. Notably, our method transforms the estimated geometry into the robot's coordinate frame, while not requiring the extrinsic parameters of the external camera to be calibrated. Our approach leverages 3D foundation models, large models pre-trained on huge datasets for 3D vision tasks, to produce initial estimates of the in-hand object. These initial estimations do not have physically correct scales and are in the camera's frame. Then, we formulate, and efficiently solve, a coordinate-alignment problem to recover accurate scales, along with a transformation of the objects to the coordinate frame of the robot. Forward kinematics mappings can subsequently be defined from the manipulator's joint angles to specified points on the object. These mappings enable the estimation of points on the held object at arbitrary configurations, enabling robot motion to be designed with respect to coordinates on the grasped objects. We empirically evaluate our approach on a robot manipulator holding a diverse set of real-world objects.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10331
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3D Foundation Models Enable Simultaneous Geometry and Pose Estimation of Grasped Objects
Zhi, Weiming
Tang, Haozhan
Zhang, Tianyi
Johnson-Roberson, Matthew
Robotics
Machine Learning
Systems and Control
Humans have the remarkable ability to use held objects as tools to interact with their environment. For this to occur, humans internally estimate how hand movements affect the object's movement. We wish to endow robots with this capability. We contribute methodology to jointly estimate the geometry and pose of objects grasped by a robot, from RGB images captured by an external camera. Notably, our method transforms the estimated geometry into the robot's coordinate frame, while not requiring the extrinsic parameters of the external camera to be calibrated. Our approach leverages 3D foundation models, large models pre-trained on huge datasets for 3D vision tasks, to produce initial estimates of the in-hand object. These initial estimations do not have physically correct scales and are in the camera's frame. Then, we formulate, and efficiently solve, a coordinate-alignment problem to recover accurate scales, along with a transformation of the objects to the coordinate frame of the robot. Forward kinematics mappings can subsequently be defined from the manipulator's joint angles to specified points on the object. These mappings enable the estimation of points on the held object at arbitrary configurations, enabling robot motion to be designed with respect to coordinates on the grasped objects. We empirically evaluate our approach on a robot manipulator holding a diverse set of real-world objects.
title 3D Foundation Models Enable Simultaneous Geometry and Pose Estimation of Grasped Objects
topic Robotics
Machine Learning
Systems and Control
url https://arxiv.org/abs/2407.10331