Real3D: Scaling Up Large Reconstruction Models with Real-World Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Hanwen, Huang, Qixing, Pavlakos, Georgios
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929383595311104
author Jiang, Hanwen
Huang, Qixing
Pavlakos, Georgios
author_facet Jiang, Hanwen
Huang, Qixing
Pavlakos, Georgios
contents The default strategy for training single-view Large Reconstruction Models (LRMs) follows the fully supervised route using large-scale datasets of synthetic 3D assets or multi-view captures. Although these resources simplify the training procedure, they are hard to scale up beyond the existing datasets and they are not necessarily representative of the real distribution of object shapes. To address these limitations, in this paper, we introduce Real3D, the first LRM system that can be trained using single-view real-world images. Real3D introduces a novel self-training framework that can benefit from both the existing synthetic data and diverse single-view real images. We propose two unsupervised losses that allow us to supervise LRMs at the pixel- and semantic-level, even for training examples without ground-truth 3D or novel views. To further improve performance and scale up the image data, we develop an automatic data curation approach to collect high-quality examples from in-the-wild images. Our experiments show that Real3D consistently outperforms prior work in four diverse evaluation settings that include real and synthetic data, as well as both in-domain and out-of-domain shapes. Code and model can be found here: https://hwjiang1510.github.io/Real3D/
format Preprint
id arxiv_https___arxiv_org_abs_2406_08479
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Real3D: Scaling Up Large Reconstruction Models with Real-World Images
Jiang, Hanwen
Huang, Qixing
Pavlakos, Georgios
Computer Vision and Pattern Recognition
The default strategy for training single-view Large Reconstruction Models (LRMs) follows the fully supervised route using large-scale datasets of synthetic 3D assets or multi-view captures. Although these resources simplify the training procedure, they are hard to scale up beyond the existing datasets and they are not necessarily representative of the real distribution of object shapes. To address these limitations, in this paper, we introduce Real3D, the first LRM system that can be trained using single-view real-world images. Real3D introduces a novel self-training framework that can benefit from both the existing synthetic data and diverse single-view real images. We propose two unsupervised losses that allow us to supervise LRMs at the pixel- and semantic-level, even for training examples without ground-truth 3D or novel views. To further improve performance and scale up the image data, we develop an automatic data curation approach to collect high-quality examples from in-the-wild images. Our experiments show that Real3D consistently outperforms prior work in four diverse evaluation settings that include real and synthetic data, as well as both in-domain and out-of-domain shapes. Code and model can be found here: https://hwjiang1510.github.io/Real3D/
title Real3D: Scaling Up Large Reconstruction Models with Real-World Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.08479