Turbo3D: Ultra-fast Text-to-3D Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866913599292702720 |
|---|---|
| author | Hu, Hanzhe Yin, Tianwei Luan, Fujun Hu, Yiwei Tan, Hao Xu, Zexiang Bi, Sai Tulsiani, Shubham Zhang, Kai |
| author_facet | Hu, Hanzhe Yin, Tianwei Luan, Fujun Hu, Yiwei Tan, Hao Xu, Zexiang Bi, Sai Tulsiani, Shubham Zhang, Kai |
| contents | We present Turbo3D, an ultra-fast text-to-3D system capable of generating high-quality Gaussian splatting assets in under one second. Turbo3D employs a rapid 4-step, 4-view diffusion generator and an efficient feed-forward Gaussian reconstructor, both operating in latent space. The 4-step, 4-view generator is a student model distilled through a novel Dual-Teacher approach, which encourages the student to learn view consistency from a multi-view teacher and photo-realism from a single-view teacher. By shifting the Gaussian reconstructor's inputs from pixel space to latent space, we eliminate the extra image decoding time and halve the transformer sequence length for maximum efficiency. Our method demonstrates superior 3D generation results compared to previous baselines, while operating in a fraction of their runtime. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_04470 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Turbo3D: Ultra-fast Text-to-3D Generation Hu, Hanzhe Yin, Tianwei Luan, Fujun Hu, Yiwei Tan, Hao Xu, Zexiang Bi, Sai Tulsiani, Shubham Zhang, Kai Computer Vision and Pattern Recognition We present Turbo3D, an ultra-fast text-to-3D system capable of generating high-quality Gaussian splatting assets in under one second. Turbo3D employs a rapid 4-step, 4-view diffusion generator and an efficient feed-forward Gaussian reconstructor, both operating in latent space. The 4-step, 4-view generator is a student model distilled through a novel Dual-Teacher approach, which encourages the student to learn view consistency from a multi-view teacher and photo-realism from a single-view teacher. By shifting the Gaussian reconstructor's inputs from pixel space to latent space, we eliminate the extra image decoding time and halve the transformer sequence length for maximum efficiency. Our method demonstrates superior 3D generation results compared to previous baselines, while operating in a fraction of their runtime. |
| title | Turbo3D: Ultra-fast Text-to-3D Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.04470 |