Structured 3D Latents for Scalable and Versatile 3D Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiang, Jianfeng, Lv, Zelong, Xu, Sicheng, Deng, Yu, Wang, Ruicheng, Zhang, Bowen, Chen, Dong, Tong, Xin, Yang, Jiaolong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916768326352896
author Xiang, Jianfeng
Lv, Zelong
Xu, Sicheng
Deng, Yu
Wang, Ruicheng
Zhang, Bowen
Chen, Dong
Tong, Xin
Yang, Jiaolong
author_facet Xiang, Jianfeng
Lv, Zelong
Xu, Sicheng
Deng, Yu
Wang, Ruicheng
Zhang, Bowen
Chen, Dong
Tong, Xin
Yang, Jiaolong
contents We introduce a novel 3D generation method for versatile and high-quality 3D asset creation. The cornerstone is a unified Structured LATent (SLAT) representation which allows decoding to different output formats, such as Radiance Fields, 3D Gaussians, and meshes. This is achieved by integrating a sparsely-populated 3D grid with dense multiview visual features extracted from a powerful vision foundation model, comprehensively capturing both structural (geometry) and textural (appearance) information while maintaining flexibility during decoding. We employ rectified flow transformers tailored for SLAT as our 3D generation models and train models with up to 2 billion parameters on a large 3D asset dataset of 500K diverse objects. Our model generates high-quality results with text or image conditions, significantly surpassing existing methods, including recent ones at similar scales. We showcase flexible output format selection and local 3D editing capabilities which were not offered by previous models. Code, model, and data will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2412_01506
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Structured 3D Latents for Scalable and Versatile 3D Generation
Xiang, Jianfeng
Lv, Zelong
Xu, Sicheng
Deng, Yu
Wang, Ruicheng
Zhang, Bowen
Chen, Dong
Tong, Xin
Yang, Jiaolong
Computer Vision and Pattern Recognition
We introduce a novel 3D generation method for versatile and high-quality 3D asset creation. The cornerstone is a unified Structured LATent (SLAT) representation which allows decoding to different output formats, such as Radiance Fields, 3D Gaussians, and meshes. This is achieved by integrating a sparsely-populated 3D grid with dense multiview visual features extracted from a powerful vision foundation model, comprehensively capturing both structural (geometry) and textural (appearance) information while maintaining flexibility during decoding. We employ rectified flow transformers tailored for SLAT as our 3D generation models and train models with up to 2 billion parameters on a large 3D asset dataset of 500K diverse objects. Our model generates high-quality results with text or image conditions, significantly surpassing existing methods, including recent ones at similar scales. We showcase flexible output format selection and local 3D editing capabilities which were not offered by previous models. Code, model, and data will be released.
title Structured 3D Latents for Scalable and Versatile 3D Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.01506