RoCap: A Robotic Data Collection Pipeline for the Pose Estimation of Appearance-Changing Objects

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Jiahao Nick, Chong, Toby, Zhou, Zhongyi, Yoshida, Hironori, Yatani, Koji, Chen, Xiang 'Anthony', Igarashi, Takeo
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913426761056256
author Li, Jiahao Nick
Chong, Toby
Zhou, Zhongyi
Yoshida, Hironori
Yatani, Koji
Chen, Xiang 'Anthony'
Igarashi, Takeo
author_facet Li, Jiahao Nick
Chong, Toby
Zhou, Zhongyi
Yoshida, Hironori
Yatani, Koji
Chen, Xiang 'Anthony'
Igarashi, Takeo
contents Object pose estimation plays a vital role in mixed-reality interactions when users manipulate tangible objects as controllers. Traditional vision-based object pose estimation methods leverage 3D reconstruction to synthesize training data. However, these methods are designed for static objects with diffuse colors and do not work well for objects that change their appearance during manipulation, such as deformable objects like plush toys, transparent objects like chemical flasks, reflective objects like metal pitchers, and articulated objects like scissors. To address this limitation, we propose Rocap, a robotic pipeline that emulates human manipulation of target objects while generating data labeled with ground truth pose information. The user first gives the target object to a robotic arm, and the system captures many pictures of the object in various 6D configurations. The system trains a model by using captured images and their ground truth pose information automatically calculated from the joint angles of the robotic arm. We showcase pose estimation for appearance-changing objects by training simple deep-learning models using the collected data and comparing the results with a model trained with synthetic data based on 3D reconstruction via quantitative and qualitative evaluation. The findings underscore the promising capabilities of Rocap.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08081
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RoCap: A Robotic Data Collection Pipeline for the Pose Estimation of Appearance-Changing Objects
Li, Jiahao Nick
Chong, Toby
Zhou, Zhongyi
Yoshida, Hironori
Yatani, Koji
Chen, Xiang 'Anthony'
Igarashi, Takeo
Robotics
Human-Computer Interaction
Object pose estimation plays a vital role in mixed-reality interactions when users manipulate tangible objects as controllers. Traditional vision-based object pose estimation methods leverage 3D reconstruction to synthesize training data. However, these methods are designed for static objects with diffuse colors and do not work well for objects that change their appearance during manipulation, such as deformable objects like plush toys, transparent objects like chemical flasks, reflective objects like metal pitchers, and articulated objects like scissors. To address this limitation, we propose Rocap, a robotic pipeline that emulates human manipulation of target objects while generating data labeled with ground truth pose information. The user first gives the target object to a robotic arm, and the system captures many pictures of the object in various 6D configurations. The system trains a model by using captured images and their ground truth pose information automatically calculated from the joint angles of the robotic arm. We showcase pose estimation for appearance-changing objects by training simple deep-learning models using the collected data and comparing the results with a model trained with synthetic data based on 3D reconstruction via quantitative and qualitative evaluation. The findings underscore the promising capabilities of Rocap.
title RoCap: A Robotic Data Collection Pipeline for the Pose Estimation of Appearance-Changing Objects
topic Robotics
Human-Computer Interaction
url https://arxiv.org/abs/2407.08081