KeyPointDiffuser: Unsupervised 3D Keypoint Learning via Latent Diffusion Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Newbury, Rhys, Zhang, Juyan, Tran, Tin, Kurniawati, Hanna, Kulić, Dana
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915651000467456
author Newbury, Rhys
Zhang, Juyan
Tran, Tin
Kurniawati, Hanna
Kulić, Dana
author_facet Newbury, Rhys
Zhang, Juyan
Tran, Tin
Kurniawati, Hanna
Kulić, Dana
contents Understanding and representing the structure of 3D objects in an unsupervised manner remains a core challenge in computer vision and graphics. Most existing unsupervised keypoint methods are not designed for unconditional generative settings, restricting their use in modern 3D generative pipelines; our formulation explicitly bridges this gap. We present an unsupervised framework for learning spatially structured 3D keypoints from point cloud data. These keypoints serve as a compact and interpretable representation that conditions an Elucidated Diffusion Model (EDM) to reconstruct the full shape. The learned keypoints exhibit repeatable spatial structure across object instances and support smooth interpolation in keypoint space, indicating that they capture geometric variation. Our method achieves strong performance across diverse object categories, yielding a 6 percentage-point improvement in keypoint consistency compared to prior approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03450
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KeyPointDiffuser: Unsupervised 3D Keypoint Learning via Latent Diffusion Models
Newbury, Rhys
Zhang, Juyan
Tran, Tin
Kurniawati, Hanna
Kulić, Dana
Computer Vision and Pattern Recognition
Machine Learning
Understanding and representing the structure of 3D objects in an unsupervised manner remains a core challenge in computer vision and graphics. Most existing unsupervised keypoint methods are not designed for unconditional generative settings, restricting their use in modern 3D generative pipelines; our formulation explicitly bridges this gap. We present an unsupervised framework for learning spatially structured 3D keypoints from point cloud data. These keypoints serve as a compact and interpretable representation that conditions an Elucidated Diffusion Model (EDM) to reconstruct the full shape. The learned keypoints exhibit repeatable spatial structure across object instances and support smooth interpolation in keypoint space, indicating that they capture geometric variation. Our method achieves strong performance across diverse object categories, yielding a 6 percentage-point improvement in keypoint consistency compared to prior approaches.
title KeyPointDiffuser: Unsupervised 3D Keypoint Learning via Latent Diffusion Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2512.03450