CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Xiao, Zhu, Minghao, Dang, Ronghao, Zhou, Guangliang, Shu, Shaolong, Lin, Feng, Liu, Chengju, Chen, Qijun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909190389235712
author Lin, Xiao
Zhu, Minghao
Dang, Ronghao
Zhou, Guangliang
Shu, Shaolong
Lin, Feng
Liu, Chengju
Chen, Qijun
author_facet Lin, Xiao
Zhu, Minghao
Dang, Ronghao
Zhou, Guangliang
Shu, Shaolong
Lin, Feng
Liu, Chengju
Chen, Qijun
contents Most of existing category-level object pose estimation methods devote to learning the object category information from point cloud modality. However, the scale of 3D datasets is limited due to the high cost of 3D data collection and annotation. Consequently, the category features extracted from these limited point cloud samples may not be comprehensive. This motivates us to investigate whether we can draw on knowledge of other modalities to obtain category information. Inspired by this motivation, we propose CLIPose, a novel 6D pose framework that employs the pre-trained vision-language model to develop better learning of object category information, which can fully leverage abundant semantic knowledge in image and text modalities. To make the 3D encoder learn category-specific features more efficiently, we align representations of three modalities in feature space via multi-modal contrastive learning. In addition to exploiting the pre-trained knowledge of the CLIP's model, we also expect it to be more sensitive with pose parameters. Therefore, we introduce a prompt tuning approach to fine-tune image encoder while we incorporate rotations and translations information in the text descriptions. CLIPose achieves state-of-the-art performance on two mainstream benchmark datasets, REAL275 and CAMERA25, and runs in real-time during inference (40FPS).
format Preprint
id arxiv_https___arxiv_org_abs_2402_15726
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge
Lin, Xiao
Zhu, Minghao
Dang, Ronghao
Zhou, Guangliang
Shu, Shaolong
Lin, Feng
Liu, Chengju
Chen, Qijun
Computer Vision and Pattern Recognition
Most of existing category-level object pose estimation methods devote to learning the object category information from point cloud modality. However, the scale of 3D datasets is limited due to the high cost of 3D data collection and annotation. Consequently, the category features extracted from these limited point cloud samples may not be comprehensive. This motivates us to investigate whether we can draw on knowledge of other modalities to obtain category information. Inspired by this motivation, we propose CLIPose, a novel 6D pose framework that employs the pre-trained vision-language model to develop better learning of object category information, which can fully leverage abundant semantic knowledge in image and text modalities. To make the 3D encoder learn category-specific features more efficiently, we align representations of three modalities in feature space via multi-modal contrastive learning. In addition to exploiting the pre-trained knowledge of the CLIP's model, we also expect it to be more sensitive with pose parameters. Therefore, we introduce a prompt tuning approach to fine-tune image encoder while we incorporate rotations and translations information in the text descriptions. CLIPose achieves state-of-the-art performance on two mainstream benchmark datasets, REAL275 and CAMERA25, and runs in real-time during inference (40FPS).
title CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.15726