G3DR: Generative 3D Reconstruction in ImageNet

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Reddy, Pradyumna, Elezi, Ismail, Deng, Jiankang
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911825368449024
author Reddy, Pradyumna
Elezi, Ismail
Deng, Jiankang
author_facet Reddy, Pradyumna
Elezi, Ismail
Deng, Jiankang
contents We introduce a novel 3D generative method, Generative 3D Reconstruction (G3DR) in ImageNet, capable of generating diverse and high-quality 3D objects from single images, addressing the limitations of existing methods. At the heart of our framework is a novel depth regularization technique that enables the generation of scenes with high-geometric fidelity. G3DR also leverages a pretrained language-vision model, such as CLIP, to enable reconstruction in novel views and improve the visual realism of generations. Additionally, G3DR designs a simple but effective sampling procedure to further improve the quality of generations. G3DR offers diverse and efficient 3D asset generation based on class or text conditioning. Despite its simplicity, G3DR is able to beat state-of-theart methods, improving over them by up to 22% in perceptual metrics and 90% in geometry scores, while needing only half of the training time. Code is available at https://github.com/preddy5/G3DR
format Preprint
id arxiv_https___arxiv_org_abs_2403_00939
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle G3DR: Generative 3D Reconstruction in ImageNet
Reddy, Pradyumna
Elezi, Ismail
Deng, Jiankang
Computer Vision and Pattern Recognition
Graphics
We introduce a novel 3D generative method, Generative 3D Reconstruction (G3DR) in ImageNet, capable of generating diverse and high-quality 3D objects from single images, addressing the limitations of existing methods. At the heart of our framework is a novel depth regularization technique that enables the generation of scenes with high-geometric fidelity. G3DR also leverages a pretrained language-vision model, such as CLIP, to enable reconstruction in novel views and improve the visual realism of generations. Additionally, G3DR designs a simple but effective sampling procedure to further improve the quality of generations. G3DR offers diverse and efficient 3D asset generation based on class or text conditioning. Despite its simplicity, G3DR is able to beat state-of-theart methods, improving over them by up to 22% in perceptual metrics and 90% in geometry scores, while needing only half of the training time. Code is available at https://github.com/preddy5/G3DR
title G3DR: Generative 3D Reconstruction in ImageNet
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2403.00939