SCOPE: Semantic Conditioning for Sim2Real Category-Level Object Pose Estimation in Robotics

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hönig, Peter, Thalhammer, Stefan, Weibel, Jean-Baptiste, Hirschmanner, Matthias, Vincze, Markus
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916976519020544
author Hönig, Peter
Thalhammer, Stefan
Weibel, Jean-Baptiste
Hirschmanner, Matthias
Vincze, Markus
author_facet Hönig, Peter
Thalhammer, Stefan
Weibel, Jean-Baptiste
Hirschmanner, Matthias
Vincze, Markus
contents Object manipulation requires accurate object pose estimation. In open environments, robots encounter unknown objects, which requires semantic understanding in order to generalize both to known categories and beyond. To resolve this challenge, we present SCOPE, a diffusion-based category-level object pose estimation model that eliminates the need for discrete category labels by leveraging DINOv2 features as continuous semantic priors. By combining these DINOv2 features with photorealistic training data and a noise model for point normals, we reduce the Sim2Real gap in category-level object pose estimation. Furthermore, injecting the continuous semantic priors via cross-attention enables SCOPE to learn canonicalized object coordinate systems across object instances beyond the distribution of known categories. SCOPE outperforms the current state of the art in synthetically trained category-level object pose estimation, achieving a relative improvement of 31.9\% on the 5$^\circ$5cm metric. Additional experiments on two instance-level datasets demonstrate generalization beyond known object categories, enabling grasping of unseen objects from unknown categories with a success rate of up to 100\%. Code available: https://github.com/hoenigpeter/scope.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SCOPE: Semantic Conditioning for Sim2Real Category-Level Object Pose Estimation in Robotics
Hönig, Peter
Thalhammer, Stefan
Weibel, Jean-Baptiste
Hirschmanner, Matthias
Vincze, Markus
Computer Vision and Pattern Recognition
Robotics
Object manipulation requires accurate object pose estimation. In open environments, robots encounter unknown objects, which requires semantic understanding in order to generalize both to known categories and beyond. To resolve this challenge, we present SCOPE, a diffusion-based category-level object pose estimation model that eliminates the need for discrete category labels by leveraging DINOv2 features as continuous semantic priors. By combining these DINOv2 features with photorealistic training data and a noise model for point normals, we reduce the Sim2Real gap in category-level object pose estimation. Furthermore, injecting the continuous semantic priors via cross-attention enables SCOPE to learn canonicalized object coordinate systems across object instances beyond the distribution of known categories. SCOPE outperforms the current state of the art in synthetically trained category-level object pose estimation, achieving a relative improvement of 31.9\% on the 5$^\circ$5cm metric. Additional experiments on two instance-level datasets demonstrate generalization beyond known object categories, enabling grasping of unseen objects from unknown categories with a success rate of up to 100\%. Code available: https://github.com/hoenigpeter/scope.
title SCOPE: Semantic Conditioning for Sim2Real Category-Level Object Pose Estimation in Robotics
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2509.24572