Yolo-Key-6D: Single Stage Monocular 6D Pose Estimation with Keypoint Enhancements

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Çetiner, Kemal Alperen, Ekenel, Hazım Kemal
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908865568702464
author Çetiner, Kemal Alperen
Ekenel, Hazım Kemal
author_facet Çetiner, Kemal Alperen
Ekenel, Hazım Kemal
contents Estimating the 6D pose of objects from a single RGB image is a critical task for robotics and extended reality applications. However, state-of-the-art multi stage methods often suffer from high latency, making them unsuitable for real time use. In this paper, we present Yolo-Key-6D, a novel single stage, end-to-end framework for monocular 6D pose estimation designed for both speed and accuracy. Our approach enhances a YOLO based architecture by integrating an auxiliary head that regresses the 2D projections of an object's 3D bounding box corners. This keypoint detection task significantly improves the network's understanding of 3D geometry. For stable end-to-end training, we directly regress rotation using a continuous 9D representation projected to SO(3) via singular value decomposition. On the LINEMOD and LINEMOD-Occluded benchmarks, YOLO-Key-6D achieves competitive accuracy scores of 96.24% and 69.41%, respectively, with the ADD(-S) 0.1d metric, while proving itself to operate in real time. Our results demonstrate that a carefully designed single stage method can provide a practical and effective balance of performance and efficiency for real world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03879
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Yolo-Key-6D: Single Stage Monocular 6D Pose Estimation with Keypoint Enhancements
Çetiner, Kemal Alperen
Ekenel, Hazım Kemal
Computer Vision and Pattern Recognition
68T45
I.4.8
Estimating the 6D pose of objects from a single RGB image is a critical task for robotics and extended reality applications. However, state-of-the-art multi stage methods often suffer from high latency, making them unsuitable for real time use. In this paper, we present Yolo-Key-6D, a novel single stage, end-to-end framework for monocular 6D pose estimation designed for both speed and accuracy. Our approach enhances a YOLO based architecture by integrating an auxiliary head that regresses the 2D projections of an object's 3D bounding box corners. This keypoint detection task significantly improves the network's understanding of 3D geometry. For stable end-to-end training, we directly regress rotation using a continuous 9D representation projected to SO(3) via singular value decomposition. On the LINEMOD and LINEMOD-Occluded benchmarks, YOLO-Key-6D achieves competitive accuracy scores of 96.24% and 69.41%, respectively, with the ADD(-S) 0.1d metric, while proving itself to operate in real time. Our results demonstrate that a carefully designed single stage method can provide a practical and effective balance of performance and efficiency for real world deployment.
title Yolo-Key-6D: Single Stage Monocular 6D Pose Estimation with Keypoint Enhancements
topic Computer Vision and Pattern Recognition
68T45
I.4.8
url https://arxiv.org/abs/2603.03879