From Words to Poses: Enhancing Novel Object Pose Estimation with Vision Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Pulli, Tessa, Thalhammer, Stefan, Schwaiger, Simon, Vincze, Markus
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914943929942016
author Pulli, Tessa
Thalhammer, Stefan
Schwaiger, Simon
Vincze, Markus
author_facet Pulli, Tessa
Thalhammer, Stefan
Schwaiger, Simon
Vincze, Markus
contents Robots are increasingly envisioned to interact in real-world scenarios, where they must continuously adapt to new situations. To detect and grasp novel objects, zero-shot pose estimators determine poses without prior knowledge. Recently, vision language models (VLMs) have shown considerable advances in robotics applications by establishing an understanding between language input and image input. In our work, we take advantage of VLMs zero-shot capabilities and translate this ability to 6D object pose estimation. We propose a novel framework for promptable zero-shot 6D object pose estimation using language embeddings. The idea is to derive a coarse location of an object based on the relevancy map of a language-embedded NeRF reconstruction and to compute the pose estimate with a point cloud registration method. Additionally, we provide an analysis of LERF's suitability for open-set object pose estimation. We examine hyperparameters, such as activation thresholds for relevancy maps and investigate the zero-shot capabilities on an instance- and category-level. Furthermore, we plan to conduct robotic grasping experiments in a real-world setting.
format Preprint
id arxiv_https___arxiv_org_abs_2409_05413
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle From Words to Poses: Enhancing Novel Object Pose Estimation with Vision Language Models
Pulli, Tessa
Thalhammer, Stefan
Schwaiger, Simon
Vincze, Markus
Computer Vision and Pattern Recognition
Robotics
Robots are increasingly envisioned to interact in real-world scenarios, where they must continuously adapt to new situations. To detect and grasp novel objects, zero-shot pose estimators determine poses without prior knowledge. Recently, vision language models (VLMs) have shown considerable advances in robotics applications by establishing an understanding between language input and image input. In our work, we take advantage of VLMs zero-shot capabilities and translate this ability to 6D object pose estimation. We propose a novel framework for promptable zero-shot 6D object pose estimation using language embeddings. The idea is to derive a coarse location of an object based on the relevancy map of a language-embedded NeRF reconstruction and to compute the pose estimate with a point cloud registration method. Additionally, we provide an analysis of LERF's suitability for open-set object pose estimation. We examine hyperparameters, such as activation thresholds for relevancy maps and investigate the zero-shot capabilities on an instance- and category-level. Furthermore, we plan to conduct robotic grasping experiments in a real-world setting.
title From Words to Poses: Enhancing Novel Object Pose Estimation with Vision Language Models
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2409.05413