PickScan: Object discovery and reconstruction from handheld interactions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: van der Brugge, Vincent, Pollefeys, Marc, Tenenbaum, Joshua B., Tewari, Ayush, Jatavallabhula, Krishna Murthy
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909394322587648
author van der Brugge, Vincent
Pollefeys, Marc
Tenenbaum, Joshua B.
Tewari, Ayush
Jatavallabhula, Krishna Murthy
author_facet van der Brugge, Vincent
Pollefeys, Marc
Tenenbaum, Joshua B.
Tewari, Ayush
Jatavallabhula, Krishna Murthy
contents Reconstructing compositional 3D representations of scenes, where each object is represented with its own 3D model, is a highly desirable capability in robotics and augmented reality. However, most existing methods rely heavily on strong appearance priors for object discovery, therefore only working on those classes of objects on which the method has been trained, or do not allow for object manipulation, which is necessary to scan objects fully and to guide object discovery in challenging scenarios. We address these limitations with a novel interaction-guided and class-agnostic method based on object displacements that allows a user to move around a scene with an RGB-D camera, hold up objects, and finally outputs one 3D model per held-up object. Our main contribution to this end is a novel approach to detecting user-object interactions and extracting the masks of manipulated objects. On a custom-captured dataset, our pipeline discovers manipulated objects with 78.3% precision at 100% recall and reconstructs them with a mean chamfer distance of 0.90cm. Compared to Co-Fusion, the only comparable interaction-based and class-agnostic baseline, this corresponds to a reduction in chamfer distance of 73% while detecting 99% fewer false positives.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11196
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PickScan: Object discovery and reconstruction from handheld interactions
van der Brugge, Vincent
Pollefeys, Marc
Tenenbaum, Joshua B.
Tewari, Ayush
Jatavallabhula, Krishna Murthy
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Robotics
I.4.5
Reconstructing compositional 3D representations of scenes, where each object is represented with its own 3D model, is a highly desirable capability in robotics and augmented reality. However, most existing methods rely heavily on strong appearance priors for object discovery, therefore only working on those classes of objects on which the method has been trained, or do not allow for object manipulation, which is necessary to scan objects fully and to guide object discovery in challenging scenarios. We address these limitations with a novel interaction-guided and class-agnostic method based on object displacements that allows a user to move around a scene with an RGB-D camera, hold up objects, and finally outputs one 3D model per held-up object. Our main contribution to this end is a novel approach to detecting user-object interactions and extracting the masks of manipulated objects. On a custom-captured dataset, our pipeline discovers manipulated objects with 78.3% precision at 100% recall and reconstructs them with a mean chamfer distance of 0.90cm. Compared to Co-Fusion, the only comparable interaction-based and class-agnostic baseline, this corresponds to a reduction in chamfer distance of 73% while detecting 99% fewer false positives.
title PickScan: Object discovery and reconstruction from handheld interactions
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
Machine Learning
Robotics
I.4.5
url https://arxiv.org/abs/2411.11196