VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kupyn, Orest, Khvedchenia, Eugene, Rupprecht, Christian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916508316205056
author Kupyn, Orest
Khvedchenia, Eugene
Rupprecht, Christian
author_facet Kupyn, Orest
Khvedchenia, Eugene
Rupprecht, Christian
contents Human head detection, keypoint estimation, and 3D head model fitting are essential tasks with many applications. However, traditional real-world datasets often suffer from bias, privacy, and ethical concerns, and they have been recorded in laboratory environments, which makes it difficult for trained models to generalize. Here, we introduce \method -- a large-scale synthetic dataset generated with diffusion models for human head detection and 3D mesh estimation. Our dataset comprises over 1 million high-resolution images, each annotated with detailed 3D head meshes, facial landmarks, and bounding boxes. Using this dataset, we introduce a new model architecture capable of simultaneous head detection and head mesh reconstruction from a single image in a single step. Through extensive experimental evaluations, we demonstrate that models trained on our synthetic data achieve strong performance on real images. Furthermore, the versatility of our dataset makes it applicable across a broad spectrum of tasks, offering a general and comprehensive representation of human heads.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18245
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
Kupyn, Orest
Khvedchenia, Eugene
Rupprecht, Christian
Computer Vision and Pattern Recognition
Machine Learning
Human head detection, keypoint estimation, and 3D head model fitting are essential tasks with many applications. However, traditional real-world datasets often suffer from bias, privacy, and ethical concerns, and they have been recorded in laboratory environments, which makes it difficult for trained models to generalize. Here, we introduce \method -- a large-scale synthetic dataset generated with diffusion models for human head detection and 3D mesh estimation. Our dataset comprises over 1 million high-resolution images, each annotated with detailed 3D head meshes, facial landmarks, and bounding boxes. Using this dataset, we introduce a new model architecture capable of simultaneous head detection and head mesh reconstruction from a single image in a single step. Through extensive experimental evaluations, we demonstrate that models trained on our synthetic data achieve strong performance on real images. Furthermore, the versatility of our dataset makes it applicable across a broad spectrum of tasks, offering a general and comprehensive representation of human heads.
title VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2407.18245