Generative 6D Pose Estimation via Conditional Flow Matching

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hamza, Amir, Boscaini, Davide, Li, Weihang, Busam, Benjamin, Poiesi, Fabio
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911462930251776
author Hamza, Amir
Boscaini, Davide
Li, Weihang
Busam, Benjamin
Poiesi, Fabio
author_facet Hamza, Amir
Boscaini, Davide
Li, Weihang
Busam, Benjamin
Poiesi, Fabio
contents Existing methods for instance-level 6D pose estimation typically rely on neural networks that either directly regress the pose in $\mathrm{SE}(3)$ or estimate it indirectly via local feature matching. The former struggle with object symmetries, while the latter fail in the absence of distinctive local features. To overcome these limitations, we propose a novel formulation of 6D pose estimation as a conditional flow matching problem in $\mathbb{R}^3$. We introduce Flose, a generative method that infers object poses via a denoising process conditioned on local features. While prior approaches based on conditional flow matching perform denoising solely based on geometric guidance, Flose integrates appearance-based semantic features to mitigate ambiguities caused by object symmetries. We further incorporate RANSAC-based registration to handle outliers. We validate Flose on five datasets from the established BOP benchmark. Flose outperforms prior methods with an average improvement of +4.5 Average Recall. Project Website : https://tev-fbk.github.io/Flose/
format Preprint
id arxiv_https___arxiv_org_abs_2602_19719
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Generative 6D Pose Estimation via Conditional Flow Matching
Hamza, Amir
Boscaini, Davide
Li, Weihang
Busam, Benjamin
Poiesi, Fabio
Computer Vision and Pattern Recognition
Existing methods for instance-level 6D pose estimation typically rely on neural networks that either directly regress the pose in $\mathrm{SE}(3)$ or estimate it indirectly via local feature matching. The former struggle with object symmetries, while the latter fail in the absence of distinctive local features. To overcome these limitations, we propose a novel formulation of 6D pose estimation as a conditional flow matching problem in $\mathbb{R}^3$. We introduce Flose, a generative method that infers object poses via a denoising process conditioned on local features. While prior approaches based on conditional flow matching perform denoising solely based on geometric guidance, Flose integrates appearance-based semantic features to mitigate ambiguities caused by object symmetries. We further incorporate RANSAC-based registration to handle outliers. We validate Flose on five datasets from the established BOP benchmark. Flose outperforms prior methods with an average improvement of +4.5 Average Recall. Project Website : https://tev-fbk.github.io/Flose/
title Generative 6D Pose Estimation via Conditional Flow Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.19719