Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jeon, Subin, Cho, In, Hong, Junyoung, Kim, Seon Joo
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912486132809728
author Jeon, Subin
Cho, In
Hong, Junyoung
Kim, Seon Joo
author_facet Jeon, Subin
Cho, In
Hong, Junyoung
Kim, Seon Joo
contents This paper introduces KeyDiff3D, a framework for unsupervised monocular 3D keypoints estimation that accurately predicts 3D keypoints from a single image. While previous methods rely on manual annotations or calibrated multi-view images, both of which are expensive to collect, our method enables monocular 3D keypoints estimation using only a collection of single-view images. To achieve this, we leverage powerful geometric priors embedded in a pretrained multi-view diffusion model. In our framework, this model generates multi-view images from a single image, serving as a supervision signal to provide 3D geometric cues to our model. We also use the diffusion model as a powerful 2D multi-view feature extractor and construct 3D feature volumes from its intermediate representations. This transforms implicit 3D priors learned by the diffusion model into explicit 3D features. Beyond accurate keypoints estimation, we further introduce a pipeline that enables manipulation of 3D objects generated by the diffusion model. Experimental results on diverse aspects and datasets, including Human3.6M, Stanford Dogs, and several in-the-wild and out-of-domain datasets, highlight the effectiveness of our method in terms of accuracy, generalization, and its ability to enable manipulation of 3D objects generated by the diffusion model from a single image.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12336
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors
Jeon, Subin
Cho, In
Hong, Junyoung
Kim, Seon Joo
Computer Vision and Pattern Recognition
This paper introduces KeyDiff3D, a framework for unsupervised monocular 3D keypoints estimation that accurately predicts 3D keypoints from a single image. While previous methods rely on manual annotations or calibrated multi-view images, both of which are expensive to collect, our method enables monocular 3D keypoints estimation using only a collection of single-view images. To achieve this, we leverage powerful geometric priors embedded in a pretrained multi-view diffusion model. In our framework, this model generates multi-view images from a single image, serving as a supervision signal to provide 3D geometric cues to our model. We also use the diffusion model as a powerful 2D multi-view feature extractor and construct 3D feature volumes from its intermediate representations. This transforms implicit 3D priors learned by the diffusion model into explicit 3D features. Beyond accurate keypoints estimation, we further introduce a pipeline that enables manipulation of 3D objects generated by the diffusion model. Experimental results on diverse aspects and datasets, including Human3.6M, Stanford Dogs, and several in-the-wild and out-of-domain datasets, highlight the effectiveness of our method in terms of accuracy, generalization, and its ability to enable manipulation of 3D objects generated by the diffusion model from a single image.
title Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.12336