DKPMV: Dense Keypoints Fusion from Multi-View RGB Frames for 6D Pose Estimation of Textureless Objects

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jiahong, Wang, Jinghao, Wang, Zi, Wang, Ziwen, Guan, Banglei, Yu, Qifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915549060005888
author Chen, Jiahong
Wang, Jinghao
Wang, Zi
Wang, Ziwen
Guan, Banglei
Yu, Qifeng
author_facet Chen, Jiahong
Wang, Jinghao
Wang, Zi
Wang, Ziwen
Guan, Banglei
Yu, Qifeng
contents 6D pose estimation of textureless objects is valuable for industrial robotic applications, yet remains challenging due to the frequent loss of depth information. Current multi-view methods either rely on depth data or insufficiently exploit multi-view geometric cues, limiting their performance. In this paper, we propose DKPMV, a pipeline that achieves dense keypoint-level fusion using only multi-view RGB images as input. We design a three-stage progressive pose optimization strategy that leverages dense multi-view keypoint geometry information. To enable effective dense keypoint fusion, we enhance the keypoint network with attentional aggregation and symmetry-aware training, improving prediction accuracy and resolving ambiguities on symmetric objects. Extensive experiments on the ROBI dataset demonstrate that DKPMV outperforms state-of-the-art multi-view RGB approaches and even surpasses the RGB-D methods in the majority of cases. The code will be available soon.
format Preprint
id arxiv_https___arxiv_org_abs_2510_10933
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DKPMV: Dense Keypoints Fusion from Multi-View RGB Frames for 6D Pose Estimation of Textureless Objects
Chen, Jiahong
Wang, Jinghao
Wang, Zi
Wang, Ziwen
Guan, Banglei
Yu, Qifeng
Computer Vision and Pattern Recognition
Robotics
6D pose estimation of textureless objects is valuable for industrial robotic applications, yet remains challenging due to the frequent loss of depth information. Current multi-view methods either rely on depth data or insufficiently exploit multi-view geometric cues, limiting their performance. In this paper, we propose DKPMV, a pipeline that achieves dense keypoint-level fusion using only multi-view RGB images as input. We design a three-stage progressive pose optimization strategy that leverages dense multi-view keypoint geometry information. To enable effective dense keypoint fusion, we enhance the keypoint network with attentional aggregation and symmetry-aware training, improving prediction accuracy and resolving ambiguities on symmetric objects. Extensive experiments on the ROBI dataset demonstrate that DKPMV outperforms state-of-the-art multi-view RGB approaches and even surpasses the RGB-D methods in the majority of cases. The code will be available soon.
title DKPMV: Dense Keypoints Fusion from Multi-View RGB Frames for 6D Pose Estimation of Textureless Objects
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2510.10933