NovelGS: Consistent Novel-view Denoising via Large Gaussian Reconstruction Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jinpeng, Xu, Jiale, Cheng, Weihao, Gao, Yiming, Wang, Xintao, Shan, Ying, Tang, Yansong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910714110672896
author Liu, Jinpeng
Xu, Jiale
Cheng, Weihao
Gao, Yiming
Wang, Xintao
Shan, Ying
Tang, Yansong
author_facet Liu, Jinpeng
Xu, Jiale
Cheng, Weihao
Gao, Yiming
Wang, Xintao
Shan, Ying
Tang, Yansong
contents We introduce NovelGS, a diffusion model for Gaussian Splatting (GS) given sparse-view images. Recent works leverage feed-forward networks to generate pixel-aligned Gaussians, which could be fast rendered. Unfortunately, the method was unable to produce satisfactory results for areas not covered by the input images due to the formulation of these methods. In contrast, we leverage the novel view denoising through a transformer-based network to generate 3D Gaussians. Specifically, by incorporating both conditional views and noisy target views, the network predicts pixel-aligned Gaussians for each view. During training, the rendered target and some additional views of the Gaussians are supervised. During inference, the target views are iteratively rendered and denoised from pure noise. Our approach demonstrates state-of-the-art performance in addressing the multi-view image reconstruction challenge. Due to generative modeling of unseen regions, NovelGS effectively reconstructs 3D objects with consistent and sharp textures. Experimental results on publicly available datasets indicate that NovelGS substantially surpasses existing image-to-3D frameworks, both qualitatively and quantitatively. We also demonstrate the potential of NovelGS in generative tasks, such as text-to-3D and image-to-3D, by integrating it with existing multiview diffusion models. We will make the code publicly accessible.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16779
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NovelGS: Consistent Novel-view Denoising via Large Gaussian Reconstruction Model
Liu, Jinpeng
Xu, Jiale
Cheng, Weihao
Gao, Yiming
Wang, Xintao
Shan, Ying
Tang, Yansong
Computer Vision and Pattern Recognition
We introduce NovelGS, a diffusion model for Gaussian Splatting (GS) given sparse-view images. Recent works leverage feed-forward networks to generate pixel-aligned Gaussians, which could be fast rendered. Unfortunately, the method was unable to produce satisfactory results for areas not covered by the input images due to the formulation of these methods. In contrast, we leverage the novel view denoising through a transformer-based network to generate 3D Gaussians. Specifically, by incorporating both conditional views and noisy target views, the network predicts pixel-aligned Gaussians for each view. During training, the rendered target and some additional views of the Gaussians are supervised. During inference, the target views are iteratively rendered and denoised from pure noise. Our approach demonstrates state-of-the-art performance in addressing the multi-view image reconstruction challenge. Due to generative modeling of unseen regions, NovelGS effectively reconstructs 3D objects with consistent and sharp textures. Experimental results on publicly available datasets indicate that NovelGS substantially surpasses existing image-to-3D frameworks, both qualitatively and quantitatively. We also demonstrate the potential of NovelGS in generative tasks, such as text-to-3D and image-to-3D, by integrating it with existing multiview diffusion models. We will make the code publicly accessible.
title NovelGS: Consistent Novel-view Denoising via Large Gaussian Reconstruction Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.16779