Zero123-6D: Zero-shot Novel View Synthesis for RGB Category-level 6D Pose Estimation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Di Felice, Francesco, Remus, Alberto, Gasperini, Stefano, Busam, Benjamin, Ott, Lionel, Tombari, Federico, Siegwart, Roland, Avizzano, Carlo Alberto
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910546304958464
author Di Felice, Francesco
Remus, Alberto
Gasperini, Stefano
Busam, Benjamin
Ott, Lionel
Tombari, Federico
Siegwart, Roland
Avizzano, Carlo Alberto
author_facet Di Felice, Francesco
Remus, Alberto
Gasperini, Stefano
Busam, Benjamin
Ott, Lionel
Tombari, Federico
Siegwart, Roland
Avizzano, Carlo Alberto
contents Estimating the pose of objects through vision is essential to make robotic platforms interact with the environment. Yet, it presents many challenges, often related to the lack of flexibility and generalizability of state-of-the-art solutions. Diffusion models are a cutting-edge neural architecture transforming 2D and 3D computer vision, outlining remarkable performances in zero-shot novel-view synthesis. Such a use case is particularly intriguing for reconstructing 3D objects. However, localizing objects in unstructured environments is rather unexplored. To this end, this work presents Zero123-6D, the first work to demonstrate the utility of Diffusion Model-based novel-view-synthesizers in enhancing RGB 6D pose estimation at category-level, by integrating them with feature extraction techniques. Novel View Synthesis allows to obtain a coarse pose that is refined through an online optimization method introduced in this work to deal with intra-category geometric differences. In such a way, the outlined method shows reduction in data requirements, removal of the necessity of depth information in zero-shot category-level 6D pose estimation task, and increased performance, quantitatively demonstrated through experiments on the CO3D dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14279
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zero123-6D: Zero-shot Novel View Synthesis for RGB Category-level 6D Pose Estimation
Di Felice, Francesco
Remus, Alberto
Gasperini, Stefano
Busam, Benjamin
Ott, Lionel
Tombari, Federico
Siegwart, Roland
Avizzano, Carlo Alberto
Computer Vision and Pattern Recognition
Estimating the pose of objects through vision is essential to make robotic platforms interact with the environment. Yet, it presents many challenges, often related to the lack of flexibility and generalizability of state-of-the-art solutions. Diffusion models are a cutting-edge neural architecture transforming 2D and 3D computer vision, outlining remarkable performances in zero-shot novel-view synthesis. Such a use case is particularly intriguing for reconstructing 3D objects. However, localizing objects in unstructured environments is rather unexplored. To this end, this work presents Zero123-6D, the first work to demonstrate the utility of Diffusion Model-based novel-view-synthesizers in enhancing RGB 6D pose estimation at category-level, by integrating them with feature extraction techniques. Novel View Synthesis allows to obtain a coarse pose that is refined through an online optimization method introduced in this work to deal with intra-category geometric differences. In such a way, the outlined method shows reduction in data requirements, removal of the necessity of depth information in zero-shot category-level 6D pose estimation task, and increased performance, quantitatively demonstrated through experiments on the CO3D dataset.
title Zero123-6D: Zero-shot Novel View Synthesis for RGB Category-level 6D Pose Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.14279