LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yabo, Yang, Chen, Fang, Jiemin, Zhang, Xiaopeng, Xie, Lingxi, Shen, Wei, Dai, Wenrui, Xiong, Hongkai, Tian, Qi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913609728131072
author Chen, Yabo
Yang, Chen
Fang, Jiemin
Zhang, Xiaopeng
Xie, Lingxi
Shen, Wei
Dai, Wenrui
Xiong, Hongkai
Tian, Qi
author_facet Chen, Yabo
Yang, Chen
Fang, Jiemin
Zhang, Xiaopeng
Xie, Lingxi
Shen, Wei
Dai, Wenrui
Xiong, Hongkai
Tian, Qi
contents Single-image 3D reconstruction remains a fundamental challenge in computer vision due to inherent geometric ambiguities and limited viewpoint information. Recent advances in Latent Video Diffusion Models (LVDMs) offer promising 3D priors learned from large-scale video data. However, leveraging these priors effectively faces three key challenges: (1) degradation in quality across large camera motions, (2) difficulties in achieving precise camera control, and (3) geometric distortions inherent to the diffusion process that damage 3D consistency. We address these challenges by proposing LiftImage3D, a framework that effectively releases LVDMs' generative priors while ensuring 3D consistency. Specifically, we design an articulated trajectory strategy to generate video frames, which decomposes video sequences with large camera motions into ones with controllable small motions. Then we use robust neural matching models, i.e. MASt3R, to calibrate the camera poses of generated frames and produce corresponding point clouds. Finally, we propose a distortion-aware 3D Gaussian splatting representation, which can learn independent distortions between frames and output undistorted canonical Gaussians. Extensive experiments demonstrate that LiftImage3D achieves state-of-the-art performance on two challenging datasets, i.e. LLFF, DL3DV, and Tanks and Temples, and generalizes well to diverse in-the-wild images, from cartoon illustrations to complex real-world scenes.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09597
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
Chen, Yabo
Yang, Chen
Fang, Jiemin
Zhang, Xiaopeng
Xie, Lingxi
Shen, Wei
Dai, Wenrui
Xiong, Hongkai
Tian, Qi
Computer Vision and Pattern Recognition
Graphics
Single-image 3D reconstruction remains a fundamental challenge in computer vision due to inherent geometric ambiguities and limited viewpoint information. Recent advances in Latent Video Diffusion Models (LVDMs) offer promising 3D priors learned from large-scale video data. However, leveraging these priors effectively faces three key challenges: (1) degradation in quality across large camera motions, (2) difficulties in achieving precise camera control, and (3) geometric distortions inherent to the diffusion process that damage 3D consistency. We address these challenges by proposing LiftImage3D, a framework that effectively releases LVDMs' generative priors while ensuring 3D consistency. Specifically, we design an articulated trajectory strategy to generate video frames, which decomposes video sequences with large camera motions into ones with controllable small motions. Then we use robust neural matching models, i.e. MASt3R, to calibrate the camera poses of generated frames and produce corresponding point clouds. Finally, we propose a distortion-aware 3D Gaussian splatting representation, which can learn independent distortions between frames and output undistorted canonical Gaussians. Extensive experiments demonstrate that LiftImage3D achieves state-of-the-art performance on two challenging datasets, i.e. LLFF, DL3DV, and Tanks and Temples, and generalizes well to diverse in-the-wild images, from cartoon illustrations to complex real-world scenes.
title LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2412.09597