Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Yujing, Sun, Caiyi, Liu, Yuan, Ma, Yuexin, Yiu, Siu Ming
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909403323564032
author Sun, Yujing
Sun, Caiyi
Liu, Yuan
Ma, Yuexin
Yiu, Siu Ming
author_facet Sun, Yujing
Sun, Caiyi
Liu, Yuan
Ma, Yuexin
Yiu, Siu Ming
contents In this paper, we present a novel generalizable object pose estimation method to determine the object pose using only one RGB image. Unlike traditional approaches that rely on instance-level object pose estimation and necessitate extensive training data, our method offers generalization to unseen objects without extensive training, operates with a single reference image of the object, and eliminates the need for 3D object models or multiple views of the object. These characteristics are achieved by utilizing a diffusion model to generate novel-view images and conducting a two-sided matching on these generated images. Quantitative experiments demonstrate the superiority of our method over existing pose estimation techniques across both synthetic and real-world datasets. Remarkably, our approach maintains strong performance even in scenarios with significant viewpoint changes, highlighting its robustness and versatility in challenging conditions. The code will be re leased at https://github.com/scy639/Gen2SM.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15860
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching
Sun, Yujing
Sun, Caiyi
Liu, Yuan
Ma, Yuexin
Yiu, Siu Ming
Computer Vision and Pattern Recognition
In this paper, we present a novel generalizable object pose estimation method to determine the object pose using only one RGB image. Unlike traditional approaches that rely on instance-level object pose estimation and necessitate extensive training data, our method offers generalization to unseen objects without extensive training, operates with a single reference image of the object, and eliminates the need for 3D object models or multiple views of the object. These characteristics are achieved by utilizing a diffusion model to generate novel-view images and conducting a two-sided matching on these generated images. Quantitative experiments demonstrate the superiority of our method over existing pose estimation techniques across both synthetic and real-world datasets. Remarkably, our approach maintains strong performance even in scenarios with significant viewpoint changes, highlighting its robustness and versatility in challenging conditions. The code will be re leased at https://github.com/scy639/Gen2SM.
title Generalizable Single-view Object Pose Estimation by Two-side Generating and Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.15860