SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910378334617600 |
|---|---|
| author | Chen, Yamei Di, Yan Zhai, Guangyao Manhardt, Fabian Zhang, Chenyangguang Zhang, Ruida Tombari, Federico Navab, Nassir Busam, Benjamin |
| author_facet | Chen, Yamei Di, Yan Zhai, Guangyao Manhardt, Fabian Zhang, Chenyangguang Zhang, Ruida Tombari, Federico Navab, Nassir Busam, Benjamin |
| contents | Category-level object pose estimation, aiming to predict the 6D pose and 3D size of objects from known categories, typically struggles with large intra-class shape variation. Existing works utilizing mean shapes often fall short of capturing this variation. To address this issue, we present SecondPose, a novel approach integrating object-specific geometric features with semantic category priors from DINOv2. Leveraging the advantage of DINOv2 in providing SE(3)-consistent semantic features, we hierarchically extract two types of SE(3)-invariant geometric features to further encapsulate local-to-global object-specific information. These geometric features are then point-aligned with DINOv2 features to establish a consistent object representation under SE(3) transformations, facilitating the mapping from camera space to the pre-defined canonical space, thus further enhancing pose estimation. Extensive experiments on NOCS-REAL275 demonstrate that SecondPose achieves a 12.4% leap forward over the state-of-the-art. Moreover, on a more complex dataset HouseCat6D which provides photometrically challenging objects, SecondPose still surpasses other competitors by a large margin. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_11125 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation Chen, Yamei Di, Yan Zhai, Guangyao Manhardt, Fabian Zhang, Chenyangguang Zhang, Ruida Tombari, Federico Navab, Nassir Busam, Benjamin Computer Vision and Pattern Recognition Category-level object pose estimation, aiming to predict the 6D pose and 3D size of objects from known categories, typically struggles with large intra-class shape variation. Existing works utilizing mean shapes often fall short of capturing this variation. To address this issue, we present SecondPose, a novel approach integrating object-specific geometric features with semantic category priors from DINOv2. Leveraging the advantage of DINOv2 in providing SE(3)-consistent semantic features, we hierarchically extract two types of SE(3)-invariant geometric features to further encapsulate local-to-global object-specific information. These geometric features are then point-aligned with DINOv2 features to establish a consistent object representation under SE(3) transformations, facilitating the mapping from camera space to the pre-defined canonical space, thus further enhancing pose estimation. Extensive experiments on NOCS-REAL275 demonstrate that SecondPose achieves a 12.4% leap forward over the state-of-the-art. Moreover, on a more complex dataset HouseCat6D which provides photometrically challenging objects, SecondPose still surpasses other competitors by a large margin. |
| title | SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2311.11125 |