SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Simoni, Alessandro, Catalini, Riccardo, Di Nucci, Davide, Borghi, Guido, Davoli, Davide, Garattoni, Lorenzo, Francesca, Gianpiero, Kawana, Yuki, Vezzani, Roberto
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913072876093440
author Simoni, Alessandro
Catalini, Riccardo
Di Nucci, Davide
Borghi, Guido
Davoli, Davide
Garattoni, Lorenzo
Francesca, Gianpiero
Kawana, Yuki
Vezzani, Roberto
author_facet Simoni, Alessandro
Catalini, Riccardo
Di Nucci, Davide
Borghi, Guido
Davoli, Davide
Garattoni, Lorenzo
Francesca, Gianpiero
Kawana, Yuki
Vezzani, Roberto
contents Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular, these issues are caused by 2D joint locations that can be mapped to multiple 3D positions, inducing multiple possible final poses. Following these considerations, we propose leveraging diffusion-based models generation capability to predict multiple hypotheses and aggregate them in a final accurate pose. Therefore, we introduce SnapPose3D, a pose-lifting framework trained deterministically to denoise 3D poses conditioned on both visual context and 2D pose features. SnapPose3D adopts a probabilistic approach during inference, generating multiple hypotheses through random sampling from a unit Gaussian distribution. Unlike most previous methods that address pose ambiguity by processing temporal sequences, SnapPose3D uses single frames as input, avoiding tracking and limiting computational cost, data acquisition complexity, and the need for online, real-time applications. We extensively evaluate SnapPose3D on well-known benchmarks for the 3D human pose estimation task showing its ability to generate and aggregate accurate hypotheses that lead to state-of-the-art results.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26620
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses
Simoni, Alessandro
Catalini, Riccardo
Di Nucci, Davide
Borghi, Guido
Davoli, Davide
Garattoni, Lorenzo
Francesca, Gianpiero
Kawana, Yuki
Vezzani, Roberto
Computer Vision and Pattern Recognition
Depth ambiguity and joint uncertainty are the two main obstacles in obtaining accurate human pose predictions by 2D-to-3D lifting methods proposed in the literature. In particular, these issues are caused by 2D joint locations that can be mapped to multiple 3D positions, inducing multiple possible final poses. Following these considerations, we propose leveraging diffusion-based models generation capability to predict multiple hypotheses and aggregate them in a final accurate pose. Therefore, we introduce SnapPose3D, a pose-lifting framework trained deterministically to denoise 3D poses conditioned on both visual context and 2D pose features. SnapPose3D adopts a probabilistic approach during inference, generating multiple hypotheses through random sampling from a unit Gaussian distribution. Unlike most previous methods that address pose ambiguity by processing temporal sequences, SnapPose3D uses single frames as input, avoiding tracking and limiting computational cost, data acquisition complexity, and the need for online, real-time applications. We extensively evaluate SnapPose3D on well-known benchmarks for the 3D human pose estimation task showing its ability to generate and aggregate accurate hypotheses that lead to state-of-the-art results.
title SnapPose3D: Diffusion-Based Single-Frame 2D-to-3D Lifting of Human Poses
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.26620