RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Junwen, Vutukur, Shishir Reddy, Yu, Peter KT, Navab, Nassir, Ilic, Slobodan, Busam, Benjamin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912662760194048
author Huang, Junwen
Vutukur, Shishir Reddy
Yu, Peter KT
Navab, Nassir
Ilic, Slobodan
Busam, Benjamin
author_facet Huang, Junwen
Vutukur, Shishir Reddy
Yu, Peter KT
Navab, Nassir
Ilic, Slobodan
Busam, Benjamin
contents Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the correct template often leads to inaccurate pose predictions. To address this, we reformulate template-based object pose estimation as a ray alignment problem, where the viewing directions from multiple posed template images are learned to align with a non-posed query image. Inspired by recent progress in diffusion-based camera pose estimation, we embed this formulation into a diffusion transformer architecture that aligns a query image with a set of posed templates. We reparameterize object rotation using object-centered camera rays and model object translation by extending scale-invariant translation estimation to dense translation offsets. Our model leverages geometric priors from the templates to guide accurate query pose inference. A coarse-to-fine training strategy based on narrowed template sampling improves performance without modifying the network architecture. Extensive experiments across multiple benchmark datasets show competitive results of our method compared to state-of-the-art approaches in unseen object pose estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18521
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation
Huang, Junwen
Vutukur, Shishir Reddy
Yu, Peter KT
Navab, Nassir
Ilic, Slobodan
Busam, Benjamin
Computer Vision and Pattern Recognition
Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the correct template often leads to inaccurate pose predictions. To address this, we reformulate template-based object pose estimation as a ray alignment problem, where the viewing directions from multiple posed template images are learned to align with a non-posed query image. Inspired by recent progress in diffusion-based camera pose estimation, we embed this formulation into a diffusion transformer architecture that aligns a query image with a set of posed templates. We reparameterize object rotation using object-centered camera rays and model object translation by extending scale-invariant translation estimation to dense translation offsets. Our model leverages geometric priors from the templates to guide accurate query pose inference. A coarse-to-fine training strategy based on narrowed template sampling improves performance without modifying the network architecture. Extensive experiments across multiple benchmark datasets show competitive results of our method compared to state-of-the-art approaches in unseen object pose estimation.
title RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.18521