Turbo3D: Ultra-fast Text-to-3D Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hu, Hanzhe, Yin, Tianwei, Luan, Fujun, Hu, Yiwei, Tan, Hao, Xu, Zexiang, Bi, Sai, Tulsiani, Shubham, Zhang, Kai
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913599292702720
author Hu, Hanzhe
Yin, Tianwei
Luan, Fujun
Hu, Yiwei
Tan, Hao
Xu, Zexiang
Bi, Sai
Tulsiani, Shubham
Zhang, Kai
author_facet Hu, Hanzhe
Yin, Tianwei
Luan, Fujun
Hu, Yiwei
Tan, Hao
Xu, Zexiang
Bi, Sai
Tulsiani, Shubham
Zhang, Kai
contents We present Turbo3D, an ultra-fast text-to-3D system capable of generating high-quality Gaussian splatting assets in under one second. Turbo3D employs a rapid 4-step, 4-view diffusion generator and an efficient feed-forward Gaussian reconstructor, both operating in latent space. The 4-step, 4-view generator is a student model distilled through a novel Dual-Teacher approach, which encourages the student to learn view consistency from a multi-view teacher and photo-realism from a single-view teacher. By shifting the Gaussian reconstructor's inputs from pixel space to latent space, we eliminate the extra image decoding time and halve the transformer sequence length for maximum efficiency. Our method demonstrates superior 3D generation results compared to previous baselines, while operating in a fraction of their runtime.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04470
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Turbo3D: Ultra-fast Text-to-3D Generation
Hu, Hanzhe
Yin, Tianwei
Luan, Fujun
Hu, Yiwei
Tan, Hao
Xu, Zexiang
Bi, Sai
Tulsiani, Shubham
Zhang, Kai
Computer Vision and Pattern Recognition
We present Turbo3D, an ultra-fast text-to-3D system capable of generating high-quality Gaussian splatting assets in under one second. Turbo3D employs a rapid 4-step, 4-view diffusion generator and an efficient feed-forward Gaussian reconstructor, both operating in latent space. The 4-step, 4-view generator is a student model distilled through a novel Dual-Teacher approach, which encourages the student to learn view consistency from a multi-view teacher and photo-realism from a single-view teacher. By shifting the Gaussian reconstructor's inputs from pixel space to latent space, we eliminate the extra image decoding time and halve the transformer sequence length for maximum efficiency. Our method demonstrates superior 3D generation results compared to previous baselines, while operating in a fraction of their runtime.
title Turbo3D: Ultra-fast Text-to-3D Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.04470