Saved in:
Bibliographic Details
Main Authors: Kuang, Liming, Velikova, Yordanka, Saleh, Mahdi, Zaech, Jan-Nico, Paudel, Danda Pani, Busam, Benjamin
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.09056
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913023366529024
author Kuang, Liming
Velikova, Yordanka
Saleh, Mahdi
Zaech, Jan-Nico
Paudel, Danda Pani
Busam, Benjamin
author_facet Kuang, Liming
Velikova, Yordanka
Saleh, Mahdi
Zaech, Jan-Nico
Paudel, Danda Pani
Busam, Benjamin
contents Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurrently, large-scale vision language models show remarkable zero-shot capabilities. In this work, we bridge these two worlds by introducing ConceptPose, a framework for object pose estimation that is both training-free and model-free. ConceptPose leverages a vision-language-model (VLM) to create open-vocabulary 3D concept maps, where each point is tagged with a concept vector derived from saliency maps. By establishing robust 3D-3D correspondences across concept maps, our approach allows precise estimation of 6DoF relative pose. Without any object or dataset-specific training, our approach achieves state-of-the-art results on common zero shot relative pose estimation benchmarks, outperforming the strongest baseline by a relative 62\% in average ADD(-S) score, including methods that utilize extensive dataset-specific training.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09056
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
Kuang, Liming
Velikova, Yordanka
Saleh, Mahdi
Zaech, Jan-Nico
Paudel, Danda Pani
Busam, Benjamin
Computer Vision and Pattern Recognition
Object pose estimation is a fundamental task in computer vision and robotics, yet most methods require extensive, dataset-specific training. Concurrently, large-scale vision language models show remarkable zero-shot capabilities. In this work, we bridge these two worlds by introducing ConceptPose, a framework for object pose estimation that is both training-free and model-free. ConceptPose leverages a vision-language-model (VLM) to create open-vocabulary 3D concept maps, where each point is tagged with a concept vector derived from saliency maps. By establishing robust 3D-3D correspondences across concept maps, our approach allows precise estimation of 6DoF relative pose. Without any object or dataset-specific training, our approach achieves state-of-the-art results on common zero shot relative pose estimation benchmarks, outperforming the strongest baseline by a relative 62\% in average ADD(-S) score, including methods that utilize extensive dataset-specific training.
title ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.09056