One2Any: One-Reference 6D Pose Estimation for Any Object

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Mengya, Li, Siyuan, Chhatkuli, Ajad, Truong, Prune, Van Gool, Luc, Tombari, Federico
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918012307636224
author Liu, Mengya
Li, Siyuan
Chhatkuli, Ajad
Truong, Prune
Van Gool, Luc
Tombari, Federico
author_facet Liu, Mengya
Li, Siyuan
Chhatkuli, Ajad
Truong, Prune
Van Gool, Luc
Tombari, Federico
contents 6D object pose estimation remains challenging for many applications due to dependencies on complete 3D models, multi-view images, or training limited to specific object categories. These requirements make generalization to novel objects difficult for which neither 3D models nor multi-view images may be available. To address this, we propose a novel method One2Any that estimates the relative 6-degrees of freedom (DOF) object pose using only a single reference-single query RGB-D image, without prior knowledge of its 3D model, multi-view data, or category constraints. We treat object pose estimation as an encoding-decoding process, first, we obtain a comprehensive Reference Object Pose Embedding (ROPE) that encodes an object shape, orientation, and texture from a single reference view. Using this embedding, a U-Net-based pose decoding module produces Reference Object Coordinate (ROC) for new views, enabling fast and accurate pose estimation. This simple encoding-decoding framework allows our model to be trained on any pair-wise pose data, enabling large-scale training and demonstrating great scalability. Experiments on multiple benchmark datasets demonstrate that our model generalizes well to novel objects, achieving state-of-the-art accuracy and robustness even rivaling methods that require multi-view or CAD inputs, at a fraction of compute.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04109
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle One2Any: One-Reference 6D Pose Estimation for Any Object
Liu, Mengya
Li, Siyuan
Chhatkuli, Ajad
Truong, Prune
Van Gool, Luc
Tombari, Federico
Computer Vision and Pattern Recognition
6D object pose estimation remains challenging for many applications due to dependencies on complete 3D models, multi-view images, or training limited to specific object categories. These requirements make generalization to novel objects difficult for which neither 3D models nor multi-view images may be available. To address this, we propose a novel method One2Any that estimates the relative 6-degrees of freedom (DOF) object pose using only a single reference-single query RGB-D image, without prior knowledge of its 3D model, multi-view data, or category constraints. We treat object pose estimation as an encoding-decoding process, first, we obtain a comprehensive Reference Object Pose Embedding (ROPE) that encodes an object shape, orientation, and texture from a single reference view. Using this embedding, a U-Net-based pose decoding module produces Reference Object Coordinate (ROC) for new views, enabling fast and accurate pose estimation. This simple encoding-decoding framework allows our model to be trained on any pair-wise pose data, enabling large-scale training and demonstrating great scalability. Experiments on multiple benchmark datasets demonstrate that our model generalizes well to novel objects, achieving state-of-the-art accuracy and robustness even rivaling methods that require multi-view or CAD inputs, at a fraction of compute.
title One2Any: One-Reference 6D Pose Estimation for Any Object
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.04109