Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Giebenhain, Simon, Kirschstein, Tobias, Rünz, Martin, Agapito, Lourdes, Nießner, Matthias
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913814862102528
author Giebenhain, Simon
Kirschstein, Tobias
Rünz, Martin
Agapito, Lourdes
Nießner, Matthias
author_facet Giebenhain, Simon
Kirschstein, Tobias
Rünz, Martin
Agapito, Lourdes
Nießner, Matthias
contents We address the 3D reconstruction of human faces from a single RGB image. To this end, we propose Pixel3DMM, a set of highly-generalized vision transformers which predict per-pixel geometric cues in order to constrain the optimization of a 3D morphable face model (3DMM). We exploit the latent features of the DINO foundation model, and introduce a tailored surface normal and uv-coordinate prediction head. We train our model by registering three high-quality 3D face datasets against the FLAME mesh topology, which results in a total of over 1,000 identities and 976K images. For 3D face reconstruction, we propose a FLAME fitting opitmization that solves for the 3DMM parameters from the uv-coordinate and normal estimates. To evaluate our method, we introduce a new benchmark for single-image face reconstruction, which features high diversity facial expressions, viewing angles, and ethnicities. Crucially, our benchmark is the first to evaluate both posed and neutral facial geometry. Ultimately, our method outperforms the most competitive baselines by over 15% in terms of geometric accuracy for posed facial expressions.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction
Giebenhain, Simon
Kirschstein, Tobias
Rünz, Martin
Agapito, Lourdes
Nießner, Matthias
Computer Vision and Pattern Recognition
Artificial Intelligence
We address the 3D reconstruction of human faces from a single RGB image. To this end, we propose Pixel3DMM, a set of highly-generalized vision transformers which predict per-pixel geometric cues in order to constrain the optimization of a 3D morphable face model (3DMM). We exploit the latent features of the DINO foundation model, and introduce a tailored surface normal and uv-coordinate prediction head. We train our model by registering three high-quality 3D face datasets against the FLAME mesh topology, which results in a total of over 1,000 identities and 976K images. For 3D face reconstruction, we propose a FLAME fitting opitmization that solves for the 3DMM parameters from the uv-coordinate and normal estimates. To evaluate our method, we introduce a new benchmark for single-image face reconstruction, which features high diversity facial expressions, viewing angles, and ethnicities. Crucially, our benchmark is the first to evaluate both posed and neutral facial geometry. Ultimately, our method outperforms the most competitive baselines by over 15% in terms of geometric accuracy for posed facial expressions.
title Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.00615