GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lan, Yushi, Zhou, Shangchen, Lyu, Zhaoyang, Hong, Fangzhou, Yang, Shuai, Dai, Bo, Pan, Xingang, Loy, Chen Change
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909573133107200
author Lan, Yushi
Zhou, Shangchen
Lyu, Zhaoyang
Hong, Fangzhou
Yang, Shuai
Dai, Bo
Pan, Xingang
Loy, Chen Change
author_facet Lan, Yushi
Zhou, Shangchen
Lyu, Zhaoyang
Hong, Fangzhou
Yang, Shuai
Dai, Bo
Pan, Xingang
Loy, Chen Change
contents While 3D content generation has advanced significantly, existing methods still face challenges with input formats, latent space design, and output representations. This paper introduces a novel 3D generation framework that addresses these challenges, offering scalable, high-quality 3D generation with an interactive Point Cloud-structured Latent space. Our framework employs a Variational Autoencoder (VAE) with multi-view posed RGB-D(epth)-N(ormal) renderings as input, using a unique latent space design that preserves 3D shape information, and incorporates a cascaded latent flow-based model for improved shape-texture disentanglement. The proposed method, GaussianAnything, supports multi-modal conditional 3D generation, allowing for point cloud, caption, and single image inputs. Notably, the newly proposed latent space naturally enables geometry-texture disentanglement, thus allowing 3D-aware editing. Experimental results demonstrate the effectiveness of our approach on multiple datasets, outperforming existing native 3D methods in both text- and image-conditioned 3D generation.
format Preprint
id arxiv_https___arxiv_org_abs_2411_08033
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation
Lan, Yushi
Zhou, Shangchen
Lyu, Zhaoyang
Hong, Fangzhou
Yang, Shuai
Dai, Bo
Pan, Xingang
Loy, Chen Change
Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
While 3D content generation has advanced significantly, existing methods still face challenges with input formats, latent space design, and output representations. This paper introduces a novel 3D generation framework that addresses these challenges, offering scalable, high-quality 3D generation with an interactive Point Cloud-structured Latent space. Our framework employs a Variational Autoencoder (VAE) with multi-view posed RGB-D(epth)-N(ormal) renderings as input, using a unique latent space design that preserves 3D shape information, and incorporates a cascaded latent flow-based model for improved shape-texture disentanglement. The proposed method, GaussianAnything, supports multi-modal conditional 3D generation, allowing for point cloud, caption, and single image inputs. Notably, the newly proposed latent space naturally enables geometry-texture disentanglement, thus allowing 3D-aware editing. Experimental results demonstrate the effectiveness of our approach on multiple datasets, outperforming existing native 3D methods in both text- and image-conditioned 3D generation.
title GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Graphics
url https://arxiv.org/abs/2411.08033