Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.09056 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913023366529024 |
|---|---|
| author | Kuang, Liming Velikova, Yordanka Saleh, Mahdi Zaech, Jan-Nico Paudel, Danda Pani Busam, Benjamin |
| author_facet | Kuang, Liming Velikova, Yordanka Saleh, Mahdi Zaech, Jan-Nico Paudel, Danda Pani Busam, Benjamin |
| contents | Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurrently, large-scale vision language models show remarkable zero-shot capabilities. In this work, we bridge these two worlds by introducing ConceptPose, a framework for object pose estimation that is both training-free and model-free. ConceptPose leverages a vision-language-model (VLM) to create open-vocabulary 3D concept maps, where each point is tagged with a concept vector derived from saliency maps. By establishing robust 3D-3D correspondences across concept maps, our approach allows precise estimation of 6DoF relative pose. Without any object or dataset-specific training, our approach achieves state-of-the-art results on common zero shot relative pose estimation benchmarks, outperforming the strongest baseline by a relative 62\% in average ADD(-S) score, including methods that utilize extensive dataset-specific training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_09056 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors Kuang, Liming Velikova, Yordanka Saleh, Mahdi Zaech, Jan-Nico Paudel, Danda Pani Busam, Benjamin Computer Vision and Pattern Recognition Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurrently, large-scale vision language models show remarkable zero-shot capabilities. In this work, we bridge these two worlds by introducing ConceptPose, a framework for object pose estimation that is both training-free and model-free. ConceptPose leverages a vision-language-model (VLM) to create open-vocabulary 3D concept maps, where each point is tagged with a concept vector derived from saliency maps. By establishing robust 3D-3D correspondences across concept maps, our approach allows precise estimation of 6DoF relative pose. Without any object or dataset-specific training, our approach achieves state-of-the-art results on common zero shot relative pose estimation benchmarks, outperforming the strongest baseline by a relative 62\% in average ADD(-S) score, including methods that utilize extensive dataset-specific training. |
| title | ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.09056 |